Clouddo

Agents

Why this blog exists: systems, cost, and agents - not model recaps

Clouddo Team, Developers · 24 Sept 2026

The useful part of AI engineering in 2026 is not the model card. It is the week after: when production, cost, and the harness fail together. This blog exists for that week.

Systems-cost-agents

Why this blog exists: systems, cost, and agents - not model recaps

The useful part of AI engineering in 2026 is not the model card.

It is the week after the model card: when the new ID lands in production, the cache hit rate collapses, the tool schema the agent depended on is suddenly rejected, the bill jumps 40% with no traffic change, and nobody can tell you whether the agent failed because the model got worse or because the harness never had a contract.

This blog exists for that week.

What we will not write

We will not recap launches.

Frontier labs now ship in clusters. In a single September week we got Claude Fable 5.1, Gemini 3.8 Flash, Muse Spark 1.3, and GPT-6 Astra. A few days later the price war started, cache windows moved, and at least one “cheaper” model broke the APIs agent stacks were pinned to. By the time a recap is published, the default model in your config has already changed twice.

Recaps have a short half-life and a high opportunity cost. They also train readers - and writers - to treat intelligence as a spectator sport. We are not interested in that sport.

We will also not write:

  • “X vs Y” leaderboard posts with no workload attached
  • tutorials that stop at a happy-path messages.create
  • AI-strategy essays that never touch a trace, a GPU, or an invoice
  • safety theater that is only a press release, or capability theater that pretends sandboxes are optional

If a new model matters here, it will matter because something in the system has to change: routing, caching, evals, permissions, or the unit economics of a loop.

What we will write

Three things, repeatedly, until they are boring enough to operationalize.

1. Systems

A model is a component. A product is a system.

The system includes the harness, the tools, the memory, the schema, the sandbox, the queue, the cache, the replica pinning, the eval suite, the rollback, and the human who gets paged when the agent books the wrong thing. Most failures we see attributed to “the model” are failures of one of those layers.

When OpenAI stretches prompt-cache reuse, or SageMaker ships prefix-aware routing, or GitHub Copilot goes fail-closed on local agent commands, or Vercel makes agent workspaces persist across disposable compute - that is the story. Not the parameter count.

We will write about:

  • harnesses that survive a mid-week model swap
  • prefill vs decode, KV cache locality, and why TTFT is a data-placement problem
  • traces that show why an agent acted, not just what it printed
  • evals that do not saturate the month after a lab announces AGI-adjacent scores on a public benchmark

If we cannot draw the box-and-arrow diagram, we do not understand the post yet.

2. Cost

Cost is not a finance footnote. It is an architecture constraint.

Token price cuts are the headline. Cache hit rate, prefix reuse, tool-call count, retry storms, oversized context, and “just use the frontier model for classification” are the actual bill.

A COLM 2026 result put a fine point on this: across a large set of text tasks, the best LLM and the best embedding model were nearly tied on quality - and the LLM cost up to a thousand times more. That is not an argument against frontier models. It is an argument against using them as the default hammer.

We will treat cost as a measurable property of a design:

  • $ per successful task, not $ per million tokens
  • cache-hit sensitivity, not list price
  • teacher → small student distillation when the production path does not need a 400B-class reasoner
  • the human time an agent saves, and the human time it creates in review and incident response

If a design cannot survive a 2× price cut and a 2× price hike, it is not a design. It is a demo that got lucky on the current rate card.

3. Agents

Agents are software that can take actions. That sentence should make an engineer reach for permissions before they reach for a bigger model.

The industry has already moved the interesting work into the runtime: agent registries, per-action governance, persistent sandboxes, coordinator threads with subagents, allowlists for irreversible tools. Amazon blocking a shopping agent is not a product spat. It is the threat model arriving on schedule.

We will write AgentOps as MLOps with a blast radius:

  • tools as an API surface, versioned and least-privileged
  • memory as state that can poison the next step
  • evals on traces, not on vibes
  • kill switches, dry-runs, and two-phase commits for anything that spends money, changes production, or talks to a customer

“The agent felt smarter this week” is not a release note. “Tool-error rate down 18%, irreversible actions now require confirmation, cost per resolved ticket down 31%” is.

A bias we will not hide

Most production “AI” work in 2026 should not call a frontier model.

Classification, routing, extraction, ranking, forecasting, and a surprising amount of retrieval quality are embedding, small-classifier, and tiny-specialist problems. Frontier models earn their keep on messy tool use, long-horizon planning, and the cases where the specification does not exist yet. Using them for everything is how teams burn the budget that should have paid for evals and a sandbox.

We will be wrong about some of those boundaries. When we are, we will publish the measurement that changed our minds.

How a post earns a slot here

A post goes up if it does at least one of the following:

  1. Changes a decision you can make this week (model pin, cache policy, tool permission, eval gate).
  2. Includes a number that came from a run, a bill, or a trace - not from a keynote.
  3. Leaves behind an artifact: a checklist, a diagram, a harness sketch, a decision tree.
  4. Still makes sense after the next model ID ships.

If the only contribution is “this lab released a thing,” it stays in our notes.

Who this is for

Engineers who ship: platform, ML, backend, and the growing number of people whose title is some variant of AI engineer and whose actual job is keeping a loop alive in production.

It is also for founders and tech leads who have to explain why the demo that looked cheap in March is an infrastructure problem in September.

It is not for people collecting model names. Those lists are already well maintained elsewhere.

What to expect

A small set of recurring series, so you know what you are opening:

  • Ship notes - one system, one metric, one failure
  • Cost anatomy - tokens, GPUs, cache, and human time in the same ledger
  • Agent runtime - tools, sandboxes, handoffs, evals
  • Small models first - when not to call the expensive model

Cadence will be irregular on purpose. We would rather publish one measured post than five reactions.

The next three posts in the queue:

  1. Prompt cache is the real price cut - a bill calculator, not a recap
  2. A coding-agent harness that survives a model swap
  3. Embeddings vs LLMs: reproduce the “nearly same quality, orders-of-magnitude different cost” result on one task

The point

Intelligence is getting cheaper to rent and more expensive to operate.

The people who will still have leverage in two years are not the ones who remember which lab won which week. They are the ones who can design a system that keeps working when the model changes, prove what it costs, and put a hard boundary around what an agent is allowed to touch.

That is the job. That is the blog.

All insights

Ask about this page