Recognizing Evidence, Not Conclusions

Yvette
Yvette Managing Partner
July 29, 2026 5 min read

A mutual recognition framework for frontier AI evaluation between INESIA and the German AI Safety and Security Institute

Declaration of independence

Fusion Collective does not sell products or services to any client whose AI systems it evaluates. The auditor cannot be the vendor.

This is a published policy, applied without exception.

Fusion Collective is not established in France or in Germany, holds no equity or commercial interest in any party whose systems would be evaluated under this framework, and is not a candidate for any evaluation contract arising from it. This document is offered as a neutral technical contribution, and its author has no interest in its adoption beyond seeing the arrangement built well.

That absence of interest is the reason a third party is the right author here. Neither institute can write the shared standard alone without the other reading it as one national methodology being installed as the norm.

Why now

On July 17, 2026, at the Franco-German Council of Ministers, the two governments committed to deepen institutional cooperation on AI safety and security.

The German Federal Ministry for Digital Transformation and Government Modernization and the Federal Ministry of the Interior, together with the German AI Safety and Security Institute, will engage with the Secrétariat général de la défense et de la sécurité nationale and the Direction générale des Entreprises in the context of INESIA. The stated aim includes mutual development on matters of AI safety and security.

The German institute was approved by the National Security Council on June 9, 2026, and is being built on existing structures, drawing technical expertise from the Federal Office for Information Security and the Federal Network Agency. As of this writing, it is not yet operational.

The European AI Office's Expert Forum, reporting on July 15, 2026, recommended structured coordination with peer institutions in partner jurisdictions, where shared evaluation methodologies can reduce duplication without compromising European sovereignty.

Three things follow. The commitment exists; however, the institutional architecture is not yet fixed, which means design choices are still available at low cost. And in the absence of shared evidence standard, which means that when INESIA and the German institute both evaluate the same model, nothing currently makes those two results comparable, transferable or mutually usable.

This document proposes what that standard should contain.

The design choice that determines everything

There are two ways to build mutual recognition; the choice between them isn’t technical but a choice about what the arrangement is for.

Recognition of conclusions. Each side accepts the other's verdict. Germany accepts that INESIA cleared a model and doesn’t re-examine it. Yes, on the surface this is extremely efficient, it’s also how you build a single point of failure. One methodology, applied once, produces a result that two states then treat as settled. If the methodology was wrong, both states are wrong, and the appearance of two independent confirmations makes the error harder to find rather than easier.

Recognition of evidence. Each side accepts the other's evaluation record as admissible input to its own judgment, without accepting the judgment itself. Germany can read what INESIA did, at what depth, against what threat model, with what access, and reach its own conclusion. Yes, this appears this less efficient, but it preserves two independent judgments, which is the entire reason for having two institutes.

This framework recommends recognition of evidence.

The Expert Forum frames coordination as a way to reduce duplication. And that’s the wrong objective for a safety function. In assurance, duplication is redundancy, and redundancy is the mechanism by which a wrong answer is caught. Efficiency and reliability point in opposite directions here. A framework built to eliminate duplicate work will, right now when it matters the most, have eliminated the only check available.

Four layers

Layer 1: A common evidence schema

Both institutes agree what an evaluation record must contain, regardless of what the evaluation concluded. This is the foundation and it’s achievable within months, because it requires no agreement on methodology.

Minimum contents:

· Object. The exact model or system version, its provenance, and how it differs from any publicly released version.

· Access conditions. What level of access was granted, by whom, under what terms, whether access was revocable during the work, and whether any element of the system was withheld.

· Threat model. What the evaluation was looking for, stated before results.

· Scope and exclusions. What was assessed and, explicitly, what was not. Exclusions carry more information than inclusions and are routinely omitted.

· Method. Named, versioned, with stated known limitations and stated conditions under which the method returns unreliable results.

· Evidence. Raw outputs retained and available for re-examination, not only summarized.

· Evaluator identity and independence position. Who performed the work and their relationship, financial and personnel, to the evaluated party.

· Confidence and residual uncertainty. What the evaluation could not determine.

· Dissent. Any disagreement within the evaluating team, recorded rather than resolved into a single voice.

A record meeting this schema is readable by the other institute whether, or not, the two agree on anything else.

Layer 2: Competence equivalence

Each institute needs a basis for treating the other's evaluators as qualified, without either adopting the other's national qualification regime.

The proposed workable mechanism is a declared competence profile per evaluation, mapped to a shared domain taxonomy: offensive cyber capability, chemical and biological uplift, loss of control and autonomy, manipulation and persuasion, discrimination and fundamental rights impact, security of the model and its weights. Recognition is granted per domain, not to the institute as a whole.

This matters because the two institutes will not be symmetrical. The German institute draws on the Federal Office for Information Security and the Federal Network Agency, which points toward cybersecurity and infrastructure. INESIA sits within a national security and defense structure alongside the industrial ministry. Pretending to “symmetry” will produce recognition in domains where one side has no real capability. However, by declaring asymmetry lets each side rely on the other where reliance is warranted and no further.

Layer 3: A methodological registry

Both institutes publish, to each other at minimum, the methods they used, versioned, with known limitations. This is the mechanism that makes Layer 1 records interpretable, and it’s also the mechanism that prevents convergence by accident.

The registry should carry a deliberate diversity obligation. Where both institutes have adopted the same method for a domain, that fact is flagged, and at least one is expected to maintain an independent alternative. Convergence should be a decision, not a drift.

Layer 4: A divergence protocol

This is the layer that will be omitted and it’s the one that matters the most.

When two institutes evaluate the same model and reach different conclusions, most mutual recognition arrangements treat that as a defect to be reconciled. In safety evaluation, divergence is the highest value output the system can produce. It means one method saw something the other did not.

The protocol should specify:

  • Divergence is recorded and reported, never silently reconciled.
  • Neither institute may withdraw or amend a finding solely because the other disagrees.
  • A joint technical examination determines the source of divergence: different access, different threat model, different method sensitivity, or genuine disagreement on the same evidence. Each of those has different implications and they must not be collapsed.
  • The result of that examination is published to the AI Office and to the other's supervising ministry, even if it is not resolved.
  • Unresolved divergence is a valid and final state.

An arrangement that cannot end in recorded disagreement is not an assurance arrangement. It is a coordination arrangement wearing assurance vocabulary.

The problem nobody has raised

The two institutes sit in different very security cultures and this will break the arrangement at the very first classified finding unless it’s designed for now.

INESIA operates within a national security and defense structure.

The German institute sits with the digital and interior ministries and draws on a federal information security agency.

These carry very different clearance regimes, different classification vocabularies, different rules on what may or may not be shared with a foreign counterpart, and different default answers to whether a finding with defense implications can cross a border at all.

The framework therefore requires a classification annex agreed upon from the outset, covering:

· A mapping between the two classification systems, at working level.

· A defined minimum that must be shared even where the substance cannot be, specifically the existence of a finding, its domain and its severity band, so that neither side is unaware that the other holds something.

· Named cleared points of contact on each side.

· A rule that classification may not be used to withhold a methodological limitation. A limitation in method is not a national security fact.

That last proposed rule will most likely be the one with the most resistance and honestly, it’s the one that actually protects both institutes from each other's blind spots.

Preserving redundancy in practice

Two provisions, both relatively cheap:

Reserved re-testing. A fixed proportion of evaluations, agreed annually, is independently re-run by the other institute using a declared different method, with the second institute blind to the first's conclusion until it files its own. This is the most reliable way to learn whether the arrangement is producing agreement because the models are being read correctly or because the methods share an assumption.

Blind sequencing. Where both institutes evaluate the same model, the second doesn’t receive the first's conclusion until its own threat model and scope are fixed and recorded.

Extension path

This should be designed from the start as a template, not as a bilateral instrument.

The Expert Forum recommends pooling with a coalition of trusted partners including the United Kingdom, Canada, Japan, the Republic of Korea, Australia, New Zealand and India, and urges the European Union to convene an inaugural summit of such a coalition. The Commission's Cybersecurity and AI Action Plan points to the international network coordinated by the UK AI Security Institute as the platform for developing common evaluation methodologies.

A Franco-German arrangement built as a bilateral special case will have to be rebuilt to join either. However, built as a template, with the evidence schema and the divergence protocol as the portable elements, it now becomes the reference the coalition adopts. And that’s the difference between a cooperation agreement and a standard.

Governance

A neutral secretariat. Rotating, or held by a 3rd party with no evaluation role, maintaining the registry, the schema version history and the divergence record. Neither institute should hold the record of its own disagreements.

Separation of accreditation from commissioning. Neither institute should qualify the external evaluators whose findings it also commissions and relies on. Where external evaluators are used, qualification should rest with national accreditation bodies operating a common scheme. This mirrors the recommendation in the companion paper (Who Evaluates the Evaluators) published on evaluator independence and resolves a conflict the Expert Forum report leaves open, since the same institutions are asked both to secure access to frontier models and to independently judge them.

Annual public reporting. Number of evaluations, domains covered, divergences recorded and their disposition. Not findings, which will often be restricted. The metadata alone would tell European legislators more about the state of frontier AI oversight than anything currently published.

What this is not

It is not a conformity assessment scheme and creates no presumption of conformity under the AI Act. It doesn’t assume or presume to substitute for the AI Office's supervisory and enforcement powers over general-purpose AI models under Articles 91 to 93, which take effect from August 2, 2026. It does not qualify evaluators, nor does it decide whether a model may be placed on the Union market. It does one thing and one thing only: it makes the work of two national institutes legible to each other, and to the Commission, without collapsing them into one.

Sources

· Joint Statement by the Federal Republic of Germany and the French Republic on Cooperation in AI Safety and Security, and Federal Ministry press release 43/2026, July 17, 2026.

· German National Security Council decision to establish an AI safety institute, June 9, 2026, drawing on the Federal Office for Information Security and the Federal Network Agency.

· European AI Office, Enhancing competitiveness, sovereignty and security of the European Union in frontier AI, Publications Office of the European Union, 2026, pp. 23 to 25 (sovereign audit, evaluation and verification capacity; three tracks; coordination with peer institutions; coalition of trusted partners) and p. 20 (verification layer).

· European Commission, Action Plan on Cybersecurity and Artificial Intelligence, COM(2026) 577 final, July 7, 2026, https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX%3A52026DC0577

· Regulation (EU) 2024/1689 (AI Act), Articles 91 to 93.

This document reflects the author's professional assessment. It does not constitute legal advice.

Related Articles