Skip to content
All comparisons
Metered cloud agent services

Krazimo vs Microsoft Foundry

Microsoft charges nothing to run the agent and meters the tools it calls. The difference worth your time is what "limit answers to your data" actually does.

The service itself is free

Stated plainly by Microsoft. We are not going to price it as a subscription.

A setting is not a mechanism

Their own docs say the model attempts to rely on your documents.

Refusal in the tool

An unlisted path fails before any text is generated.

Microsoft says the quiet part in a full sentence

On the Foundry Agent Service pricing page: There is no additional charge for creating or running Foundry-native agents using prompts and workflows.1 You pay for model tokens and for the tools an agent invokes — knowledge connections, search, code interpreter sessions, hosted compute. Anyone comparing against Azure by adding up subscriptions is comparing against something that does not exist.

The sentence worth reading twice

Azure OpenAI On Your Data has a setting for limiting answers to your documents. Microsoft's documentation describes what it does with admirable precision: When this setting is enabled, the model attempts to only rely on your documents for responses.2 Read that as an engineer rather than as a buyer. It is a description of an intention — the model attempts — and it is competing with the user's phrasing, the retrieved material, the model's priors and the previous eleven turns. It usually wins. Usually is a strange word to find at the centre of a compliance story, and Microsoft deserves credit for wording it honestly rather than as a guarantee.

What a mechanism looks like instead

Our agent can only open a document by naming a path, and the tool that opens paths will only open one that appeared in a listing it was given. Ask for anything else and the call fails — not the model declining, the call failing, in code with no context window and no opinion. That is why declined is a value of a field here rather than a phrase to detect in prose, and why the behaviour does not change with the model, the temperature or how the question was asked.

The arithmetic

On price, they win, and for the same reason Bedrock does

The agent service is free and the meter sits on the tools. Below a certain volume that is cheaper than any flat plan, ours included. The honest comparison is not a number, it is a question: what does your bill do in the month your traffic triples, and what do you have at the end of it that a compliance owner can sign?

Depending on who is asking
Platform engineering

“We are a Microsoft shop.”

Then the tooling and the identity story are already yours, and that is worth a lot. What is not in the box is retrieval that structurally cannot answer off-library.

Compliance

“Is "limit responses" a control?”

It is a prompt-level instruction, per Microsoft's own wording. Treat it as a strong default, not as an enforcement, and ask what happens on the run where it does not hold.

Product

“How do we route an escalation?”

On a decision field carried by every response — answered, declined or escalated — rather than by pattern-matching apologetic prose.

What Microsoft Foundry Agent Service publishes

Every line here is a figure from their own page, with the wording it was read from in the footnotes. Nothing in this table is our characterisation of their product.

 Metered cloud agent services
The agent service itselfNo additional charge1
Limiting answers to your documentsA setting — the model attempts to rely on your documents2
Where the meter sits

The same turn, billed two ways. Every step happens either way — the difference is whether each one carries its own charge, or the turn does.

QuestionScreenRetrieveModelGuardsAnswerMETERED PER TOOL CALLRetrieve is billed per callbilledGuards is billed per callbilledThe agent service is free; the knowledge and tool calls inside the turn are what bills.one runOne turn, one run — retrieval and guards included, not itemised.
How Krazimo Studio is built

Our own architecture, stated as such. Whether it is better than the alternative is your call to make; these are the facts you would be deciding on.

  • Retrieval navigates a named index — no embedding step, and no vector store to run or re-index
  • A path no listing offered is refused by the tool itself, before any text is generated
  • Guards are deterministic rules on the finished answer, each reported by name with its reason
  • One version pins the library, the policy, the evaluation suite and the model, and can be restored
  • Usage is returned on the call that incurred it — the same number that bills
  • Unlimited agents on every plan, and no share of your model spend
Questions we get

Is the agent service really free?

Microsoft states it on their own pricing page, quoted and linked below. You pay for model tokens, tools, knowledge connections and hosted compute.

Is "limit responses to your data" enough?

It is a good default and it holds most of the time. Whether most of the time is enough depends entirely on what your agent is answering questions about.

Can we keep using Azure models?

Yes. Point your workspace at your own provider key and inference bills to you directly, at whatever rate you have negotiated.

Sources

Every external figure on this page, with the wording it was read from and the date it was read. Competitor prices change; if one of these is out of date, tell us and it gets corrected here rather than argued about.

  1. 1Microsoft, Foundry Agent Service pricing

    “There is no additional charge for creating or running Foundry-native agents using prompts and workflows.”

  2. 2Microsoft, Azure OpenAI On Your Data

    “When this setting is enabled, the model attempts to only rely on your documents for responses.”

Every number on this page has a source. That is the product, in a different medium.

$99 a month to start. Unlimited agents, no share of your model spend, and a record a compliance owner can sign.