Anthropic has launched a dedicated site, claude.dev, for teams building products on Claude and Claude Code. It features articles, videos, and build logs from the company's own engineers: guides on selecting evals, distributing model "effort," estimating task costs on Opus 5.5, and case studies on optimizing claude.ai. This is not a model release or an API change, but a consolidation of the vendor's scattered engineering materials into a single official primary resource.

image
image
image

What happened

Anthropic has revamped its approach to engineering publications and launched a dedicated site, claude.dev, aimed at developers creating products on Claude and Claude Code. It features articles, videos, and build logs from Anthropic's own engineers. The materials focus on practical questions: how to automate eval design and hillclimbing (iterative improvement of results based on a chosen metric); how to distribute model "effort" in Claude Code; how much a task costs on Opus 5.5; and how claude.ai was sped up 3x in two weeks. Publications are released regularly: on September 25, 2026, the article "Building with Claude Sonnet 5.5" was published, and on September 28, 2026, "Automating eval design and hillclimbing." The site also features video sessions on how companies Ramp and DoorDash use Claude Code.

Context

Previously, engineering guides for Claude and Claude Code were scattered across separate blogs and channels, and teams building on these models often relied on third-party retellings. With the launch of claude.dev, the model vendor began sharing agent engineering practices — evals, skills, context engineering — directly through its own primary channel. This is notable against the broader backdrop: eval methodology is an area where vendors rarely disclose internal details, so publishing such materials from the model's author stands out. In essence, this is a platform distribution step, not an independent content initiative: the vendor is investing in ensuring developers build on its stack consciously and remain within its ecosystem.

Why this matters for the industry

For the industry, the value lies not in the announcement itself, but in the methodology. Production teams on Claude Code get a single primary source instead of third-party retellings, immediately reducing reliance on secondary interpretations of vendor practices. The official task cost benchmark for Opus 5.5 provides a basis for recalculating the unit economics of agentic systems, and discrepancies between teams' own measurements and published figures become a signal for re-verification. If the publication cadence is maintained — two fresh materials were released in late September 2026 — claude.dev could become the de facto standard channel for sharing agent engineering practices for the Anthropic stack, and vendor competition will continue to shift from the "bare model" to the ecosystem: first-class documentation, evals, skills, and first-hand build logs.

Why this matters for users

Those already developing on Claude Code or considering agentic pipelines should bookmark the site as a reference. It can be applied immediately without new API access: compare your own eval process with the "Automating eval design and hillclimbing" guide, review model "effort" distribution in Claude Code, recalculate the cost of your agentic tasks using the vendor's Opus 5.5 price benchmark, and watch the Ramp and DoorDash video sessions for transferable integration solutions. However, it is important to understand the boundaries: the launch does not bring changes to pricing, limits, or interfaces; it is a knowledge-sharing channel, not the emergence of new model capabilities — a browser bookmark is more useful here than waiting for updates.

What is still unknown / limitations

The provided materials show section and post titles, but not their content: the metrics, baselines, datasets, and reproducibility conditions from the "Automating eval design and hillclimbing" guide are unknown, so the title cannot be considered a ready-made transferable methodology. Claims about a 3x speedup of claude.ai in two weeks and the cost of a task on Opus 5.5 are given without measurement methodology and, until details are published, are more in a marketing register. It is unclear what exactly the "effort" parameter in Claude Code represents and how its effect is measured. The thesis that the new site will replace scattered blogs does not follow from the primary sources — this is the author's interpretation of the original post. Long-term scenarios, such as the consolidation of agent engineering standards around vendor documentation, remain speculative: there is no reliable basis for such a forecast in the input data.

Sources

Author

Look at AI, editorial team