The cheapest call can still produce an expensive month.
Wagtail set out to use GLM 5.3 Flash for September. The team reports 2 billion tokens across the month. The target model handled the first half and cost $68. A separate MCP prototype used 450 million tokens and cost $150 almost overnight.[1]
The prototype produced a working demo. The problem was route control. A different model and a little more effort could have reduced the reported cost, according to the same account. The lesson is narrower than "model bad" or "agents expensive." Prototype work needs its own ceiling and receipt.
A monthly model policy needs an exception log. Otherwise the exception becomes the bill.
Provider health belongs beside model choice.
Wagtail also reports degraded GLM 5.3 Flash performance during the trial. The team switched to DeepSeek V4.1 Flash and Qwen 3.8 Flash. Its provider comparison already treated model breadth, model lifecycle, energy reporting, cache hit rates, and sovereignty as operating criteria.[2]
A fallback is not a panic button if you define it before the outage. Record the trigger, substitute model, provider, task class, and result. Without that record, a temporary reroute can become an unreviewed default.
Vendor efficiency is an input, not your fuel log.
Z.AI describes GLM 5.3 Flash as a 320-billion-parameter model with 18 billion activated parameters. Its documentation claims lower attention computation and key-value cache use than GLM 5.3. It also advertises a one-million-token context window and tool calling.[3]
Those specifications help form a shortlist. They do not measure your cache reuse, failed calls, review time, or useful output. Keep vendor claims in one column and local observations in another.
Count outcomes beside tokens.
Wagtail's operating recommendations call for token, cache, and cost reporting. They warn that cost estimates can drift because rate data changes. The same page recommends spend quotas and a multi-provider coding tool so the route can change when required.[4]
Add one result field for each expensive run. Name the accepted patch, shipped prototype, closed issue, or rejected attempt. A month with fewer tokens and no useful result is not efficient. A month with a costly prototype may be worth it, but the owner should be able to point to the artifact.
Run the next trial with four records.
- Set a monthly spend ceiling and a separate prototype ceiling.
- Record the requested model, served model, provider, cache data, cost, and failure state for each run.
- Define the fallback trigger and keep a matched task for comparing the default and fallback routes.
- Review useful outcomes, failed attempts, and exceptions before choosing the next month's default.
The fuel book below drafts that review. It does not call a model, read provider billing, measure energy, or verify a result.