Services
Every engagement scores your system against the same eight dimensions, on the same four levels. What changes is coverage and depth.
Every level, depth and readiness threshold named on this page is defined by the Ordoia Assurance Levels, v1.0, published 2026-09-19 at ordoia.com/oal/v1.0.
Where to start
Coverage is how many dimensions. Depth is how far we go on each. They are two separate axes, and the price follows both.
| Coverage | Inspected | Tested | Sustained |
|---|---|---|---|
| Four dimensions | Audit£2,500 · 1 week | not offered | not offered |
| ↓ baseline top-up · £2,500 · ~1 week | audit plus top-up reaches exactly the same place as a baseline taken directly | ||
| All eight dimensions | Baseline£5,000 · 2 weeks | Reviewfrom £9,000 · 3 weeks | Retainerfrom £3,000/month · 6-month minimum |
The cells marked not offered are not gaps waiting to be filled. Tested and sustained depth are only sold across all eight dimensions, because a partial adversarial pass would produce a scorecard whose silences a reader could not interpret.
There is no wrong entry point and no penalty for starting small. The audit (£2,500) plus a later top-up of the remaining four dimensions (£2,500) reaches exactly the same place as a baseline taken directly (£5,000). Every assessment can be followed by any other.
1 · Agent grounding and readiness audit
One week · £2,500 fixed · four dimensions · inspected depth
The fastest way to find out whether your agent is telling the truth.
We read the code and configuration, run a probe set against your system, and score four of the eight dimensions: grounding and entailment, tool-use integrity, evaluation discipline, and observability and failure detection.
The remaining four — authorisation and data boundary, refusal and instruction-boundary robustness, model and upgrade control, and execution bounds and cost attribution — appear on your scorecard as not assessed. They can't be established honestly in a week without adversarial testing, and we would rather name the gap than imply coverage. You can add them later at inspected depth for £2,500, or take them at tested depth as part of a review.
Method note: we've built grounded retrieval against live public registers, including Companies House, using a verified-or-null discipline — a record is returned only when it can be traced to a source, never because it is plausible.
You receive: a four-dimension scored assessment against our published rubric, a ranked defect register with reproduction steps, and a deterministic eval harness you own and keep.
The harness is yours to run and extend. We don't score you with it — our findings come from a separate probe set that stays with us, because an assessor who grades your system using the instrument they handed you is grading their own work.
2 · Agent production readiness review
Three weeks · from £9,000, fixed before we start · eight dimensions · tested depth
All eight dimensions, established adversarially rather than by inspection.
Paraphrase sets run against your refusal paths. Deny-case scoring against real identities — testing that the system refuses correctly, not only that it answers correctly. Fault injection against every tool: forced errors, timeouts and malformed returns, with the user-visible behaviour asserted for each. Upgrade candidates compared against your eval set before promotion. Per-turn token economics and data egress mapped against your actual use cases.
This is the depth at which a readiness decision — internal use, client-facing, auditor-facing — can actually be evidenced rather than asserted.
Method note: our evaluation harness pattern scores the distribution across repeated runs rather than a single sample, because a test that passes once tells you nothing about a non-deterministic system.
You receive: the full eight-dimension assessment at tested depth, a defect register with evidence, a prioritised remediation plan your own engineers can execute, a written readiness position against each threshold, and an eval suite wired for CI.
Reports are addressed to you. If you need to put an assessment in front of a third party — your own client, an investor, a regulator — tell us at scoping, so it can be scoped and priced properly rather than assumed.
3 · Agent reliability retainer
From £3,000/month · six-month minimum · eight dimensions · sustained depth
For systems already in front of users. The question is no longer what level you're at — it's whether you're still there.
Each month we re-establish your scores and report six measures: level movement with cause; eval-set currency, including cases added from real traffic and incidents and cases retired as obsolete; upgrade events and what moved when your vendor shipped; alert hygiene, including stale rules that no longer describe the system; detection lead — defects found by monitoring or regression run versus defects reported by your users; and open defect aging by severity.
Detection lead is the number to watch. It answers, in one ratio, whether the assurance is working.
The retainer needs a baseline to trend against. It follows a review, or opens with a baseline month — all eight dimensions at inspected depth, £5,000, delivered as month one. Take a review within three months of a baseline and we credit half the baseline fee against it.
You receive: a monthly reliability report you can put in front of a risk committee, and the tenure disclosed on the face of every one — how long we've been engaged, because a long relationship is a thing a reader should be able to weigh.
How we work
Fixed scope, fixed price, named deliverables, and a written go/no-go at the end of every engagement. Every engagement is led end-to-end by a named senior architect — no handoff to a delivery team you never met.
We don't sell day rates. We don't take greenfield "build us an agent" work where nobody owns the data governance. We don't price by volume of code.
On doing this yourselves
Your engineers can build this, and the good ones do. A team that knows the business can write a better eval set than an outsider can — they know what a correct answer looks like. Point an agent at your own repository and it will find missing mechanisms too: the fallback that lives in a prompt, the index built without an identity filter, the model version floating on an alias.
None of that produces an assessment anyone else can rely on.
What we sell is not the finding. It is that the finding came from somewhere other than the team that built the system, under a method that was fixed before we looked, retained in working papers, and signed by a named person who stands behind it. Your board, your client's procurement team and your regulator all apply the same discount to self-assessment, and they apply it for the same reason companies who can read their own accounts still pay to have them audited.
So the honest division is this. Build the harness — you will build a better one than we would, and we would rather assess a system that has one. Bring us in for the part you structurally cannot do for yourself.
Start
Audit 4 dimensions 1 week £2,500Or a 30-minute scoping call, at no charge, to establish whether the audit is the right entry point for your system. Read the rubric first — most people who read it will never contact us, and that is the intention.