Frontier pricing is not a compliance strategy
Compliance agents are token furnaces: evidence in, framework references in, analysis out. We run the ISMS Copilot API on GLM 5.2 with curated framework knowledge injected at inference, at $2.80/$8.80 per million tokens and a $0.50/$2.00 bulk lane. Here is the pricing math against the Claude price sheet, and the evidence for why a non-frontier model holds up at compliance work.

A compliance agent is a token furnace. To do its job it consumes evidence (policies, control descriptions, audit artifacts, ticket exports), consumes reference material (the framework itself: controls, clauses, mappings), and produces structured analysis. The ratio is lopsided: tens of thousands of tokens in for every thousand out, and the input side grows with every document you attach. At agent scale, where one human review becomes hundreds of automated passes, the price per token is not a line item. It is the architecture constraint. It decides whether you process the whole evidence set or a sample, whether the agent re-reads the framework for every finding or works from memory, and whether the whole idea ships at all.
So when we priced the ISMS Copilot API, we started from the workload, not from the frontier. This post is the math and the reasoning, with sources, so you can redo it yourself when prices move. They will move. Anthropic's current sheet is published here; ours is in the console. Everything below was observed on 2026-08-26.
The price sheet, side by side
Anthropic's published per-million-token rates, next to ours:
| Model | Input | Output | Same workload, blended |
|---|---|---|---|
ISMS Copilot mini (isms-mini) | $0.50 | $2.00 | $0.88 |
| Claude Haiku 4.5 | $1.00 | $5.00 | $1.88 |
| Claude Sonnet 5 | $2.00 | $10.00 | $4.00 |
ISMS Copilot GLM tier (isms-fast, isms-thinking, either region) | $2.80 | $8.80 | $4.30 |
| Claude Sonnet 4.6 | $3.00 | $15.00 | $6.00 |
| Claude Opus (4.5 through 5) | $5.00 | $25.00 | $10.00 |
The blended column uses an illustrative compliance mix, 75% input and 25% output, because compliance work reads far more than it writes. Your mix will differ; the list prices are the stable part of the comparison.
Read honestly, the table says three things. Our bulk lane is less than half of Haiku for the same work, which is the lane that matters when you are formatting or classifying ten thousand evidence artifacts. Our standard tier sits well under Sonnet 4.6 and Opus, and roughly at Sonnet 5's new list price. And if your instinct is "Sonnet 5 got cheap," hold that thought, because the list price is not the whole story on the newer Claude models.
The fine print that moves the math
Three structural details, all from Anthropic's own pricing documentation, none of which are hidden but all of which are easy to miss when you multiply list prices by token counts.
The newer tokenizer. From Anthropic's docs: "Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text." More tokens for the same text is a cost multiplier wearing a technical name. On the models where it applies (which includes Sonnet 5 and current Opus), the same policy document costs roughly 30% more tokens than the list-price multiplication suggests. Our table above does not even include that effect; on the newer models, add it.
Pinned geography costs extra there, and nothing here. Anthropic charges a 1.1x multiplier on all token categories when you pin inference to the United States with inference_geo: "us". Some compliance programs want pinned geography, and that is a legitimate requirement. On our API, the EU processing path runs on Mistral's EU-bound inference host, with zero retention of request content, at the same price as the global path. Pinning is a routing choice here, not a premium tier.
Their mitigations are real, and so is the corpus. Two fair points in Anthropic's favor, because a comparison that omits them is not a comparison. The Batch API halves input and output costs for non-time-sensitive work, and prompt caching makes repeated context cheap on subsequent calls. If your compliance agent's shape fits batch windows and cache-friendly repetition, include those in your math. But note what no discount fixes: on a raw model API, grounded compliance output requires the reference material in the prompt, and that corpus is yours to license, assemble, maintain, and keep current. Frameworks get amended, transpositions get enacted, editions get superseded. On our API, the knowledge is the included part: 101 curated framework modules today, injected on detection or by pin, billed as ordinary input tokens, disclosed in every response.
Cheap is not the argument
A cheap model that invents control numbers is expensive in the only currency that matters at scale: review time. The reason our pricing can sit where it sits is that the compliance competence does not come from the base model alone. It comes from what is in the prompt.
The mechanism, stated precisely because the imprecise version is how marketing lies: the API is not a fine-tuned model, and the knowledge is not in the weights. When your request names a framework (or you pin one), the server injects the curated reference module for that framework into the prompt, alongside a published system prompt, before the model speaks. The response tells you exactly which modules ran, via the x-isms-frameworks header and a disclosure object with a token estimate. Grounding you can verify beats recall you cannot see.
The model behind the aliases is GLM 5.2, a strong open-weights model with a 1M-token context. We picked it through a recorded evaluation, not a vibe: a blind-judged, 36-generation compliance eval where GLM 5.2 with injected knowledge scored 87.8 (fast) and 90.0 (thinking) against Mistral Small's 73.3 and 74.4 on the same tasks with the same injection; a citation-trap suite where GLM passed 5 of 5 traps designed to catch confident invention of compliance facts; and a pre-registered gate for the published system prompt (20/20 accuracy, 24/24 refusal of intellectual-property elicitation, with the failed first attempt recorded rather than buried). The full record, with sample sizes and caveats, is on the new model quality and grounding docs page. We quote it with its limits: these are our internal evals, N is small, and we claim no quality superiority over Claude, because no head-to-head benchmark against raw frontier models exists yet. One is in design now, and its premise is to publish losses alongside wins.
What we can say is narrower and stronger: when a module is selected, the model answers from a maintained, verification-stamped reference instead of training recall, and you can see which reference served each response. At compliance work, that is the difference between an answer you spot-check and an answer you can put in a workpaper.
Not an EU product
A word on scope, because the compliance API conversation tends to collapse into a GDPR-and-done framing. The catalog is global, with the United States as the largest single cluster: SOC 2, HIPAA, CCPA, CMMC, FedRAMP, PCI DSS, CIS Controls, and nine NIST publications including 800-53 and 800-171. Then the United Kingdom (six modules), Japan (four), Canada (five), India (five), Switzerland (FADP, the ICT minimum standard, NCS, FINMA 23/01), Australia (Essential Eight and the Privacy Act, with AU ISM and CPS 234 landing), New Zealand, Singapore, Germany (BSI Grundschutz, C5, TISAX), France, and the 25 national transpositions of NIS 2. The live count comes from GET /v1/frameworks, which is authoritative and public with your key; it said 101 today, and it moves as frameworks get added and amended.
The EU processing path is one routing choice among two, priced identically, for programs that need it. It is not the product's identity.
When a frontier API is still the right call
Honesty cuts both ways. If your workload is genuinely frontier-hard reasoning (novel legal interpretation with no framework to anchor on, massive multi-step synthesis), a frontier model may be worth frontier prices for that step. If you need multimodal input or native tool calling, we are text-in/text-out, and the honest pattern today is a sub-agent step, not a replacement. Nothing here forbids a hybrid: frontier model for the hard synthesis, this API for everything grounded and everything at volume. The architecture question is which API handles the 95% of tokens that are reference lookups, evidence reads, and structured formatting. That is the part where frontier pricing quietly becomes unsustainable.
Redo this math yourself
Prices move; conclusions should survive the movement. Current sources: Anthropic's pricing page, our console pricing. The structural facts (tokenizer behavior, geography multipliers, knowledge included or self-assembled) move slower than prices do, and those are the ones worth arguing about.
If you are building a compliance agent or a compliance app: get an API key, point your OpenAI-compatible client at api.ismscopilot.com/v1, and read the framework knowledge and model quality docs to see exactly what lands in your prompt. The first request takes minutes. The pricing argument takes as long as you want to spend on it.
Related Posts

The silent zero: when a missing judge score becomes a measurement
In our 2026-07-17 ablation of GRC document-generation strategies (GLM-5.2 and Claude Opus 4.8 as writers and as 1-10 rubric judges), pass 2 had 79 of 446 recorded judgment rows without an overall score. The aggregator counted each one as 0. The same zero-filling had manufactured a 3.25-point effect in pass 1, and when it vanished we wrote a judge-bias theory to explain why. The paired offset between the two judges on the same documents was 0.12.

Why we do not deduplicate compliance facts by embedding similarity
In our memory dedup eval, contradictions averaged 0.938 mean cosine to their nearest stored fact and true duplicates averaged 0.940, and no single threshold cleanly separated the pairs that must not merge from the ones that should. So cosine never makes the merge decision in our pipeline.

The symmetry test: shipping a prompt fix you cannot reproduce
When a production bug will not reproduce offline, you cannot validate a prompt fix by making it succeed. You validate it by checking whether your change moves the output in a consistent direction or just adds to the model's own run-to-run noise. Here is the test we ran on 126 blind A/B pairs.
