Who Evaluates the Evaluators?

Independence and competence criteria for third-party evaluation of frontier AI models.

Yvette
Yvette Managing Partner
July 27, 2026 5 min read

Declaration of independence Fusion Collective does not sell products or services to any client whose AI systems it evaluates. The auditor cannot be the vendor.

This is a published statement that we apply without exception.

It’s not a case-by-case judgment and it’s not waived for large engagements.

Fusion Collective holds no equity, no contingent fee arrangement and no commercial interest in any party whose systems it assesses. It holds no interest in the outcome of any evaluation, and no interest in whether the criteria proposed in this paper are adopted. The criteria set out below are the standard by which we expect Fusion Collective to be judged, not only the standard proposed for others.

The gap

In July 2026 the European AI Office published the findings of its Expert Forum on Frontier AI. On page 25 the report sets out how Europe should build evaluation capacity. The first of three tracks reads: dedicated capacity for pre-deployment evaluation, third-party audit and ongoing monitoring of frontier models provided to European markets, delivered through a mix of a

accredited independent evaluators and public institutional capacity.

Two words carry that entire sentence, and neither is defined.

The word "accredited" appears once in the report. No scheme is named. No accrediting body is identified. No criteria are proposed. The word "independent" appears throughout and is never given content.

The Commission's Action Plan on Cybersecurity and Artificial Intelligence, published on July 7, 2026, commits to proposing criteria for 3rd-party evaluators for the purposes of the GPAI Code of Practice. Those criteria, as of this writing, do not yet exist. The AI Office convened a workshop on July 15, 2026, to gather expert views on exactly this question and this paper proposes content for those two words.

It does so from a specific position.

Europe isn’t the first sector to need independent assessment of things too complex for the buyer to check. Pressure vessels, medical devices, financial statements, aircraft and laboratory results all sit inside mature conformity assessment frameworks that solved the impartiality problem decades ago. And yes, we concede that those frameworks are not perfect; however, they are considerably further along than anything currently proposed for AI, coupled with the fact that they were built after the failures, not before. Europe has the perfect opportunity to build this one before.

Why this matters more than it appears

The argument for independence criteria is usually made as a matter of principle. However, the stronger argument is empirical. Self-assessments don’t work, and there’s now a documented record.

The pattern is consistent across jurisdictions. A recent national framework for agentic AI governance ran to 53 pages and recommended testing, monitoring and internal review throughout. Every case study in it described an organization governing its own systems. Independent third-party attestation appeared nowhere. The framework's acknowledgements listed 60+ vendors and agencies with no civil society, labor or affected-community participation. The result? A national document in which the party being assessed was the assessor by default.

The same structural pattern now appears in Europe. The Expert Forum that recommended accredited independent evaluators comprised European AI developers, technical researchers, businesses building on frontier AI, investors, think tanks and academia. It included no conformity assessment bodies, no national accreditation bodies and no standards organizations (per its listing of contributors). The assurance regime was designed without the assurance profession present.

The financial audit precedent is instructive and at the same time uncomfortable. Large accounting firms have audited clients while selling those same clients advisory work for decades. The conflict is documented, legislated against repeatedly and yet, still present. The lesson isn’t that independence rules fail but that independence rules written after a market has formed are written against entrenched interests and rarely survive contact with them. Europe's frontier AI evaluation market hasn’t formed yet.

The failure mode is specific. When assurance isn’t independent, the output is not a wrong answer. It’s an answer that can’t be relied on, which is even worse, because it looks like assurance and carries the reputational weight of assurance. A European evaluator whose independence is unspecified produces findings exactly as contestable as the vendor attestation it was meant to replace. Building capacity without building credibility spends all the money and doesn’t close the gap.

graded independence model is referenced in ISO/IEC 17020, Clause 4.1, and the proficiency testing requirements are addressed in ISO/IEC 17025, Clause 7.7, with the main PT-related requirement in 7.7.2.

Independence: seven dimensions

Independence is not binary. Conformity assessment practice has long graded it (ISO/IEC 17020, Clause 4.1), most intimately in the distinction between a body that is fully separate from the parties involved, a body that forms an identifiable and separated part of an organization that is itself involved, and a body that supplies or designs the item it assesses while maintaining internal separation. That graded model transfers directly. What follows below proposes what must be assessed at each grade.

  1. Commercial independence. The evaluator shouldn’t sell to the evaluated party any product or service that its own findings could recommend. This covers model development, remediation, safety tooling, red-team-as-a-service where the same engagement produces the clearance, and governance implementation. The auditor cannot be the vendor. This is the load-bearing criterion, and it is the one most commonly waived in practice, because the firms with the deepest technical capability to evaluate frontier models are frequently the firms selling into them. Take for example, Microsoft's AI Red Team documenting agents being misled by deceptive interface elements. Their Defender team identified "memory poisoning" campaigns manipulating AI assistants' memory to quietly steer future responses. Their Cyber Pulse report admits that over 80% of Fortune 500 companies are deploying AI agents, but only 47% have the necessary security controls in place. Their corporate vice president of security, Vasu Jakkal, told reporters that "agent adoption and scaling is pretty significant, but at the same time, the visibility that organizations have on the agents is very limited." This is the company now selling YOU the solution to secure the agents they're simultaneously encouraging you to deploy at scale.
  2. Financial independence. Declared limits on revenue concentration from any single evaluated party, measured across a rolling multi-year window and across affiliated entities. No equity, options, carried interest or economic exposure in the evaluated party, its competitors, or funds materially invested in either. No contingent or success-based fees. No fee structure in which a finding of non-conformity reduces future revenue.
  3. Structural independence. No common ownership, no shared ultimate parent, no board interlocks, no shared control persons. Disclosed and maintained on a public register rather than asserted at engagement.
  4. Operational independence. The evaluator determines methodology, scope and depth. The evaluated party may not restrict the evaluation to areas of its choosing, may not veto techniques, and may not terminate the engagement to prevent an adverse finding from being recorded. Where scope is negotiated, the negotiation itself must be part of the record.
  5. Reporting independence. A defined and non-waivable right to report findings to the competent authority irrespective of the evaluated party's wishes. No non-disclosure provision may override it. The evaluated party has a right of reply, not a right of suppression. Absent this, every other criterion is decorative, because an evaluator that can be silenced has no independence to protect.
  6. Personnel independence. Cooling-off periods in both directions between evaluator staff and evaluated developers, with declared thresholds by seniority and by involvement in the specific engagement. This one is going to be the most and it’s unavoidable. The Expert Forum itself identifies frontier talent as scarce, mobile and concentrated in a small number of firms. The same scarcity that makes evaluators hard to staff makes revolving-door capture near-certain if left unaddressed.
  7. Access independence. The evaluator's access to the model, its weights where applicable, its documentation and its evaluation environment must not be revocable at the evaluated party's sole discretion during an engagement, and the terms of access must be disclosed as part of the finding. An evaluation conducted under access that could be withdrawn mid-course is an evaluation conducted under pressure, and the reader is entitled to know that.

Competence: six criteria

  1. Demonstrated rather than claimed capability. The Expert Forum states the test itself, on page 18, in the context of recruiting talent: verify genuine rather than claimed expertise, for instance through a demonstrated track record at the frontier. The report never turns that test toward evaluators. And it should. Qualification should rest on completed evaluations with published methodology (i.e., proficiency testing requirements in ISO/IEC 17025, Clause 7.7), not solely on institutional prestige or self-description.
  2. Declared domain scope. No organization should hold a general qualification as an evaluator of frontier AI. Qualification should be granted per risk domain, with the domains named: offensive cyber capability, chemical and biological uplift, loss of control and autonomy, manipulation and persuasion, discrimination and fundamental rights impact, and security of the model and its weights. These require different people, different methods and different evidence. A single undifferentiated credential invites the holder to opine beyond its competence.
  3. Methodological transparency. Published, versioned methods with stated limitations and stated conditions under which the method returns unreliable results. An evaluator unwilling to publish what it does can’t be assessed by anyone, including the authority relying on its findings.
  4. Comparative calibration. Periodic participation in inter-evaluator comparison exercises, in which multiple qualified evaluators assess the same artefact independently and results are compared. Laboratory accreditation has required proficiency testing of this kind for many years, precisely because a laboratory's own confidence in its results is not evidence. No equivalent exists for AI evaluation. It should, and it is the single most practical mechanism available for distinguishing evaluators that work from evaluators that assert.
  5. Security competence. Demonstrated capability to hold model access, weights and pre-release information securely. An evaluator that leaks is a systemic risk in its own right, and the evaluators handling the most consequential models will handle the most valuable unreleased information in the industry.
  6. Continuity. Evidence of ability to sustain capability as models change. A qualification granted against a 2026 model class and never revisited certifies nothing about a 2028 one. Qualification should carry a defined validity period and surveillance obligations.

Who verifies, and the conflict that must be avoided

Three options exist. Each has a failure mode.

The AI Office designates evaluators directly. Fast, coherent, and it creates a conflict. The Commission is simultaneously negotiating access to frontier models on behalf of European institutions, as the Expert Forum recommends in 4.2.1. A body that both negotiates for continued access to a model and qualifies the evaluators who judge it holds two positions that will eventually collide.

Peer review among national institutes. Legitimate and slow, and it produces reciprocal reluctance to fail a peer.

National accreditation bodies operating a common AI evaluation scheme, with the AI Office maintaining a public register. This is the recommended route. It uses machinery that already exists across all Member States, it is Member State neutral, accreditation bodies are already required to operate impartially and not to compete with those they accredit, and it separates the accreditor from the commissioner of evaluations.

The separation principle; stated plainly. The body that commissions or relies on evaluations must not be the body that qualifies the evaluators. This is the point at which the Expert Forum's own recommendations conflict with each other, and resolving it costs nothing if it is resolved now.

Against monoculture

Page 25 of the Expert Forum report recommends structured coordination with peer institutions in partner jurisdictions, where shared evaluation methodologies can reduce duplication without compromising European sovereignty.

Reducing duplication is the wrong objective for a safety function. In assurance, duplication is redundancy, and redundancy is the mechanism by which a wrong answer is caught. If every qualified evaluator runs the same method, a shared blind spot in that method becomes a shared blind spot across the whole system, and the appearance of corroboration will make it harder to detect rather than easier.

Any qualification scheme should therefore require methodological diversity as a positive criterion, reserve a proportion of evaluations for independent re-testing by a second evaluator using a declared different method, and treat divergence between evaluators as reportable information rather than as an error to be reconciled away.

What remains unresolved

What remains are not small nor are they minor and will determine whether a qualified evaluator market actually forms.

Liability and insurance. The Expert Forum report does not mention either. Who carries the loss when a qualified evaluator clears a model that subsequently causes serious harm? Because of industry practices, it’s nearly impossible to show that the model operated that caused a harm is the exact same model that an evaluator cleared. Until this is answered, no adequately capitalized organization will accept high-stakes frontier evaluation work, and the qualification scheme will be populated by parties too small to be worth suing.

Access to models held internally. The report notes twice, at pages 12 and 24, that the most capable models are increasingly used inside developing firms before any external release, and that this narrows external oversight. A qualification scheme is meaningless for models no evaluator can reach. The Commission and ENISA are already building a European Blueprint for structured access to advanced AI capabilities, due Q4 2026 and currently scoped to cybersecurity purposes. Extending that Blueprint to cover qualified evaluator access is the most direct route available and it requires no new instrument.

Procurement as enforcement. The report identifies public procurement as a demand-side instrument for shaping the market. It does not connect procurement to evaluation. Conditioning European and Member State procurement of frontier AI on independent evaluation evidence would make the independence regime commercially self-enforcing without further legislation. This is the highest-return unused instrument in the report.

Declaration of independence, restated

Fusion Collective does not sell products or services to any client whose AI systems it evaluates. We can be your auditor. We can be your vendor. We cannot be both. This holds true for anyone, not just us.

That rule is the reason this paper exists. Every criterion proposed above is one the author accepts as binding on Fusion Collective, and readers are invited to hold the firm to it. A paper arguing for independence criteria written by a firm unwilling to meet them would be worth nothing, and the market for AI assurance will not be built by organizations that exempt themselves.

This article reflects the author's professional assessment. It does not constitute legal advice.

Related Articles