What third-party means here
It is a set of constraints we accept, not an adjective we apply to ourselves.
Nobody asks whether their tracing vendor is independent. The question only arises when someone is being asked to rely on a judgement, and it is answered by what the person making the judgement has agreed not to do. So the constraints are published, in the order they bind.
-
We do not assess systems we have built or remediated, and we do not remediate systems we have assessed. Specifying what needs building is assessment work; writing it is not ours to do, because next month we would be scoring our own code. Where we have done build work, the re-assessment belongs to someone else.
-
We take no commission, referral fee or revenue share from any tooling or model vendor. What we recommend is not what pays us.
-
Fees are fixed before we start and are never contingent on the score.
-
Every assessment names the methodology version it was performed under, and the person who performed it. Both are published. The version is printed on the face of the scorecard next to every level, because a score awarded under one version does not carry over to another.
-
Each assessment is performed by a single assessor, and there is no second reviewer on the work. No accreditation stands behind it either. What stands behind it is a published method, retained working papers, and a named person on every assessment.
-
Working papers are retained for six years. A competent assessor given the same papers should reach the same score. If they wouldn't, we've sold you an opinion.
-
We keep a conflict register, and we say no when it says no.
What the working papers contain
Every paper carries the engagement reference, the dimension, the date, the version identifier of the system examined, the assessor, and the level it supports. Alongside them, per engagement: the scope statement, the depth purchased, the methodology version, and a deviation log recording every point at which judgement was applied and why.
Reproducible in six years means a competent assessor can take the papers and reach the same levels. That requires the system's commit or deploy reference, a permanent copy of the rubric version used, the versions of every tool and scorer involved, raw outputs rather than summaries, and a manifest listing every paper so that completeness is demonstrable and a later addition is visible as one.
Six-year retention collides with client data-protection terms. We resolve it by retaining evidence in redacted or derived form — hashes, structural extracts, counts, redacted transcripts — and retaining the redaction rule alongside it, so a reader knows what was removed and how. Retaining raw client data for six years is a liability, not diligence.
The word we do not use
We say third-party rather than independent. It is the UK government's own term in the Trusted Third-Party AI Assurance Roadmap, and it is a factual statement about who is doing the looking rather than a claim about the quality of the looking. Independent is a word for a practice with apparatus behind it, and we would rather add the apparatus first.
What we cannot currently do
We do not certify, attest, accredit or approve, and no threshold outcome we issue is a verdict. A readiness threshold is reported as met or not met on the evidence obtained, which is a statement about a scorecard and one you can check yourself from the levels.
No third party may rely on a report addressed to you. If an assessment needs to go in front of your own client, an investor or a regulator, that has to be scoped and priced before we start rather than assumed afterwards.
These are the constraints of a practice at its current apparatus, stated so that you are not left to discover them.