ordoia

Your agent works in the demo. The question is what it does in front of a client.

Third-party assurance for LLM and agent systems. We assess grounding, reliability and production readiness for UK financial services, private equity and professional services firms who have already built or bought one — and now have to put it in front of clients, auditors or regulators.

Audit 4 dimensions 1 week £2,500
Grounding and entailment · one of eight dimensions
OAL 0undefined OAL 1asserted OAL 2enforced OAL 3evidenced to cross this span, answer If you deleted the sentence in your system prompt that tells the model to use only retrieved content, what in the code would stop a fabricated answer reaching a user?

A demo shows you OAL 1. A production incident shows you the difference. The gaps are drawn to difficulty, not to equal intervals — the span from asserted to enforced is the one where systems fail, and it is the same width on every dimension. The full rubric

The failures nobody demos

Agent systems don't fail loudly. They fail quietly, correctly formatted, and confidently.

A code reviewer can read your codebase and find a bug on a line. What they can't tell you is that your agent is wrong on a meaningful fraction of real queries, that retrieval quietly stopped returning the right documents last Tuesday, or that it answered correctly in March and doesn't now because your vendor shipped a new model. Those failures aren't in the code. They only appear at runtime, across many samples, and over time — which is where we look.

What that looks like in practice

  • Answers that aren't entailed by what reached the context window. Fluent, sourced-looking, and not supported by anything actually retrieved.
  • Tool calls that fail and get narrated as success. The API returned an error; the agent told the user the booking was made.
  • Retrieval that ignores who is asking. Access control is enforced at the application boundary, then bypassed by an index built without it.
  • Refusal paths that hold in testing and fold under paraphrase. The guardrail was requested in a prompt rather than enforced in code — so it holds until someone rephrases.
  • Non-determinism mistaken for a passing test. The same input passes on Tuesday and fails on Thursday. A single run tells you nothing.
  • Instructions that arrive as data and get followed anyway. A page in your knowledge base contains the sentence "ignore your previous instructions", and nothing between the index and the model tells it apart from something your user asked for.
  • Silent regression on model upgrade. Your vendor deprecates a version, behaviour shifts, and nothing in your pipeline is watching the shape of the answers.
  • Execution that runs until something else stops it. One message sends the agent round a tool loop nobody put a ceiling on, and the first report anyone reads is the invoice.

The same instrumented pass finds the ordinary defects too. In a system Ordoia built, it surfaced an agent serving fabricated company records as live results because the offline fallback was instructed in the prompt rather than enforced in code — and, separately, a signup path where every account created would have been locked out around twenty-four hours later. Neither was reported by anyone. Both were found by looking.

We find these because we instrument systems to be found out: OpenTelemetry tracing through the full agent path into Langfuse and Grafana, evaluation harnesses that run each case many times and score the distribution rather than the sample, and alerting tuned to the shape of the answers rather than uptime. Instrumentation is how we gather evidence — it is not what we sell. What we sell is the assessment it makes possible.

Score yourself in five minutes

Eight questions, one per dimension, written so that you recognise your own system. Answer them honestly and you will know roughly where you sit before anyone quotes you a price. If your answer to most of them is a sentence in a prompt, you are at OAL 1 — which is where most systems in production are, and is not a moral failing. It is a position you can now name.

The eight questions

The instrument is published

Every engagement scores your system against the same eight dimensions, on the same four levels — from a behaviour that is merely asserted in a prompt, to one enforced in code, to one evidenced by continuous verification. The rubric is published in full, free, with no email wall, because a prospect recognising their own system at OAL 1 is the whole qualification mechanism and a gate blocks precisely that recognition.

We publish no aggregate score. An OAL 0 on authorisation is not offset by an OAL 3 on cost control, and a single number would invite exactly that trade. You get a level per dimension, the depth of evidence behind each one, and the reasoning. If you need one number for a board, take the lowest.

The Ordoia Assurance Levels, v1.0

What it costs

Coverage and depth are two separate axes: how many dimensions, and how far we go on each. Prices are fixed before we start and are never contingent on the score.

Coverage × depth
Coverage Inspected Tested Sustained
Four dimensions Audit£2,500 · 1 week not offered not offered
↓ baseline top-up · £2,500 · ~1 week audit plus top-up reaches exactly the same place as a baseline taken directly
All eight dimensions Baseline£5,000 · 2 weeks Reviewfrom £9,000 · 3 weeks Retainerfrom £3,000/month · 6-month minimum

There is no wrong entry point and no penalty for starting small. Every assessment can be followed by any other. What the audit does not cover appears on your scorecard as not assessed, with the maximum level that scope could have obtained printed next to it.

What each engagement includes

What stands behind a score

Not an accreditation, and not a badge. A published method, fixed before we look and versioned so that a score can be checked against the criteria it was awarded under. Working papers retained for six years, on the standard that a competent assessor given the same papers should reach the same level. And a named person on every assessment, printed on the face of the scorecard alongside the methodology version.

Your board, your client's procurement team and your regulator all apply the same discount to self-assessment, and they apply it for the same reason companies who can read their own accounts still pay to have them audited.

What third-party means here