Accountability Architecture

Accountability Architecture

Who answers when an automated system shapes a strike?

The systems that support military decisions are becoming faster, more complex and more dependent on private suppliers, while the legal machinery for inspecting those systems, reconstructing harm and allocating responsibility remains fragmented.

The argument in 90 seconds

  1. AI already shapes military strikes — by selecting data, classifying objects, ranking targets and guiding weapons — without independently firing anything.
  2. Substantial law applies now: humanitarian law, weapons review, export control, criminal law, contract. The problem is fragmentation, not absence.
  3. A human approval is not proof of meaningful control; control depends on information, time and authority.
  4. Legal review of a product name is not review of the software, data and updates that change how it behaves.
  5. When harm occurs, responsibility scatters across state, commander, operator and supplier — and the technical evidence may not survive.
  6. Accountability is part of defence capability: procurement can build it in, before criminal law is ever needed.
How to read the labels

Every substantive claim on this site carries one of seven labels, so that current law, practice, analogy, research finding, hypothetical and proposal are never mixed:

Current law
A statement of law presently in force, as the editors understand it.
State practice
How states and militaries actually behave, supported by cited reporting or official documents.
Reported practice
Attributed reporting that has not been fully or independently verified.
Legal analogy
A precedent from another context, offered for its reasoning — not a direct authority for AI systems.
Trial Trial finding
A research finding from the simulated trial — a structured exercise, not a court judgment.
Hypothetical
A stipulated scenario used to test the framework. It asserts no fact about any real person or company.
Proposed reform — not current law
A policy recommendation. It describes what the editors argue the law should require, not what it requires now.

Where relevant, sources carry a verification status: verified, partly verified, pending legal review, or pending human check. Evidence cut-off: July 2026.

What this site is not

It is not a case against any company or state: no allegation is presented as a finding, and the one simulated case was fictionalised. It is not a claim that no law exists — the opposite. And it is not an argument against innovation: the reforms proposed are designed to reduce uncertainty for government and responsible industry alike.

The decision chain

How an automated system can shape a strike

State practice
How states and militaries actually behave, supported by cited reporting or official documents.

No fielded system needs to pull a trigger to matter legally. Decision-support software chooses which sensor data is fused, classifies what a camera sees, ranks possible targets, recommends weapon pairings, plans routes and guides munitions in their final seconds. Each function shapes what a human decides, how fast, and on what picture of the world.

Independent analysis of current practice finds exactly this: the locus of lethal decision-making is moving into software, while target recognition and navigation assistance are already routine and decision cycles compress to tens of seconds.

The chain below follows a strike from training data to investigation. Select a stage to see the decision being made, who actually controls it, the evidence that should exist, the law engaged — and where accountability typically fails.

Concepts

Decision-shaping without firing

Software can materially shape a strike through data selection, classification, ranking, recommendation, routing or terminal guidance — without independently firing anything.

The legally interesting systems are mostly not weapons in the traditional sense. They decide what a human sees, in what order, with what confidence attached, and how little time remains. Regulation aimed only at the trigger misses the influence.

SOURCES [14] Defining Autonomy: Why Software, Not Drones, Will Decide the Next War, Center for Strategic and International Studies (CSIS)

Meaningful human control

A person pressing a button is not proof of control; control depends on the information, time and authority available to make a real decision.

UK policy requires context-appropriate human involvement in weapons that identify, select and attack targets. A testable version asks: what does the reviewer see, including uncertainty? How long do they genuinely have? Can they refuse or delay? Does their behaviour collapse into rubber-stamping at realistic tempo? Each is measurable, which means each can be an acceptance condition — or ignored.

SOURCES [2] Ambitious, Safe, Responsible: our approach to the delivery of AI-enabled capability in Defence, UK Ministry of Defence[3] Proceed with Caution: Artificial Intelligence in Weapon Systems (HL Paper 16), House of Lords AI in Weapon Systems Committee

Stage · Training & operational data

The decision being made
What the system will ever be able to see: which examples, sensors, regions and labels define 'target'.
Who has practical control
Supplier data teams and, for operational feeds, the procuring force.
Evidence that should exist
Data provenance records, labelling policies, known coverage gaps, validation results by environment.
Legal / regulatory layer engaged
Contract; responsible-business due diligence; feeds Article 36 review validity.
Principal accountability failure
Provenance undocumented; a system validated on one conflict is deployed in another, and nobody can show the difference was assessed.

More than one actor and more than one legal regime can apply to the same event; each stage names the principal ones only.

Full chain reference (all stages)

1 · Training & operational data

The decision being made
What the system will ever be able to see: which examples, sensors, regions and labels define 'target'.
Who has practical control
Supplier data teams and, for operational feeds, the procuring force.
Evidence that should exist
Data provenance records, labelling policies, known coverage gaps, validation results by environment.
Legal / regulatory layer engaged
Contract; responsible-business due diligence; feeds Article 36 review validity.
Principal accountability failure
Provenance undocumented; a system validated on one conflict is deployed in another, and nobody can show the difference was assessed.

2 · Model / decision-support system

The decision being made
How observations become classifications and scores — thresholds, confidence handling, failure modes.
Who has practical control
The supplier (architecture, weights, updates); the state chooses to procure and integrate it.
Evidence that should exist
Model and version identifiers, test plans and adverse results, disclosed limitations, interface behaviour.
Legal / regulatory layer engaged
Contract and procurement law; Article 36 review of the means or method; export control if transferred.
Principal accountability failure
The deployed version cannot be identified after the fact; disclosed limitations never reached the people relying on the output.

3 · Target ranking / recommendation

The decision being made
Which candidates a human ever sees, in what order, with what confidence attached.
Who has practical control
The system in real time, within parameters set by command and the supplier's design.
Evidence that should exist
Recommendation logs: inputs, outputs, confidence values, what was displayed and what was withheld.
Legal / regulatory layer engaged
IHL precautions in attack (the humans relying on it); no regime reviews ranking behaviour directly — a gap.
Principal accountability failure
Ranking shapes the decision but is treated as 'advisory', so its influence is never separately reviewed or logged.

4 · Human review

The decision being made
Whether the recommendation becomes an intended target — the decision the law counts on being real.
Who has practical control
The operator or targeting cell, within the time and information the system and tempo allow.
Evidence that should exist
Review time per decision, information displayed, uncertainty presentation, refusal rates, training records.
Legal / regulatory layer engaged
IHL distinction, proportionality and precautions; UK policy on context-appropriate human involvement.
Principal accountability failure
Approval at interface speed: the record shows a human decision that the tempo made impossible to make.

5 · Command authorisation

The decision being made
Whether the engagement proceeds, under what rules of engagement and mission constraints.
Who has practical control
The commander, on the picture assembled upstream.
Evidence that should exist
Orders, ROE, the intelligence picture as presented at the time, legal advice given.
Legal / regulatory layer engaged
IHL; command responsibility for orders given and for failures to prevent or punish.
Principal accountability failure
The commander authorises on a machine-assembled picture nobody can later reconstruct — responsibility without inspectability.

6 · Weapon action

The decision being made
Terminal guidance and fusing — including behaviour if the link is lost or conditions change.
Who has practical control
The weapon system within its programmed envelope.
Evidence that should exist
Weapon telemetry, seeker imagery where it exists, envelope settings, abort logic.
Legal / regulatory layer engaged
IHL weapons law; the Article 36 review's assumptions about how the weapon behaves.
Principal accountability failure
Terminal behaviour diverges from reviewed assumptions and the divergence is only discoverable from logs that were optional.

7 · Incident investigation

The decision being made
What officially happened — the account every legal consequence depends on.
Who has practical control
Military investigators; potentially coroners, prosecutors, regulators, Parliament.
Evidence that should exist
Everything above, preserved, tamper-evident, and accessible to people with clearance and technical competence.
Legal / regulatory layer engaged
IHL duty to investigate grave breaches; domestic criminal law; state responsibility and remedy.
Principal accountability failure
The investigation inherits every earlier gap: missing logs, unidentifiable versions, distributed knowledge — and closes without findings.

The existing architecture

The legal architecture that exists now

Current law
A statement of law presently in force, as the editors understand it.

The problem is not an empty statute book. International humanitarian law governs the conduct of hostilities. Article 36 of Additional Protocol I requires legal review of new weapons, means and methods of warfare. Export licensing can halt transfers where risk thresholds are met. International criminal law reaches individuals; domestic law reaches companies in defined ways; contracts can require almost anything a buyer insists on.

The difficulty is that these regimes were built for different objects — a shell, a shipment, a soldier, a company — and a continuous, software-mediated decision system is all of them at once, changing monthly under commercial confidentiality.

The explorer below maps eight legal layers against the system lifecycle, from design to remedy. Each cell answers five questions: who owes the duty, what triggers it, what evidence it needs, what consequence follows, and where the gap is.

Concepts

Article 36 review

States must review whether new weapons, means or methods of warfare could be used lawfully — the question for AI is what exactly the review approved.

A review describes an object: characteristics, envelope, foreseeable effects. Machine-learning systems change with retraining, new data and updated thresholds while the product name stays constant. Unless materiality thresholds and re-review triggers are defined, the approval describes history, not the deployed system.

SOURCES [1] Additional Protocol I to the Geneva Conventions, Article 36 (new weapons), ICRC treaty database[4] Government response to the House of Lords AI in Weapon Systems Committee report, UK Government

Distinction, proportionality, precautions

IHL's targeting rules bind humans and institutions; an algorithm cannot discharge the legal judgment they owe.

Distinction requires directing attacks only at military objectives. Proportionality prohibits expected incidental harm excessive to the anticipated military advantage. Precautions require doing everything feasible to verify targets and minimise harm. A system can inform each judgment; the obligation stays with people — which is why what the system shows them, and how fast, is a legal question.

SOURCES [5] ICRC position on autonomous weapon systems, International Committee of the Red Cross

Open the framework explorer →

The core problem

Where responsibility fragments

Current law
A statement of law presently in force, as the editors understand it.

Every layer of the architecture is real. The fragmentation happens between them. The state owes IHL duties but relies on a supplier's undisclosed model. The weapons review approved version 3; version 11 is deployed. The export licence assessed hardware; behaviour lives in software updates. Criminal law needs a person with knowledge and contribution; the knowledge is distributed across a company and the contribution across a supply chain. The contract could have required logs; it did not.

No single actor designed this outcome, and no single reform fixes it. Duties are split across bodies of law with different objects, thresholds and investigators — and the technical evidence that would connect them is precisely what current arrangements fail to preserve.

That is the accountability gap: not a missing law, but missing architecture between laws.

The problem is not that no law exists. Duties are split across different bodies of law, responsibility is easily fragmented, and technical evidence may not survive.

Concepts

Traceability

The records that let anyone reconstruct, later, what a system saw, recommended and did — and which version did it.

Reconstruction needs the model version, inputs, outputs and confidence, what the operator was shown, and what they did — preserved tamper-evidently and long enough to investigate. None of it exists by default. It is a procurement requirement or it is absent, and its absence is discovered after civilians are dead.

SOURCES [13] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

What existing cases show

Nuremberg industrialist cases (IG Farben; Krupp)
Legal analogy
A precedent from another context, offered for its reasoning — not a direct authority for AI systems.

Executives were prosecuted as individuals; commercial role was no shield. But these were not convictions of the corporate entities, and the conduct — supplying in knowing furtherance of atrocity — sits far from routine defence contracting.

SOURCES [16] The Nuremberg industrialist trials (IG Farben; Krupp) — case summaries, US Holocaust Memorial Museum

Frans van Anraat (Netherlands)
Legal analogy
A precedent from another context, offered for its reasoning — not a direct authority for AI systems.

A supplier convicted of complicity in war crimes for providing a chemical-weapons precursor with proven knowledge and contribution. A supplier analogue — for a physical, single-purpose input, not continuously updated software.

SOURCES [17] Public Prosecutor v. Frans van Anraat — judgment summary, Dutch courts / ICRC national practice database

Lundin (Sweden)
Legal analogy
A precedent from another context, offered for its reasoning — not a direct authority for AI systems.

Proceedings against former executives for alleged complicity in grave crimes. Ongoing as of July 2026: cited for the existence of the theory, not its success. Status must be re-checked before publication.

SOURCES [18] Sweden v. former Lundin executives (alleged complicity in war crimes), Swedish Prosecution Authority / Stockholm District Court

Lafarge (France / US)
Legal analogy
A precedent from another context, offered for its reasoning — not a direct authority for AI systems.

Corporate liability reached through adjacent offences (terrorism financing), with separate complicity litigation. A terrorism-financing conviction is not a corporate war-crimes conviction; conflating them overstates the law.

SOURCES [19] Lafarge proceedings — terrorism-financing conviction distinguished from separate complicity litigation, French courts / US Department of Justice

None of these is a direct precedent for AI decision-support. Together they show that commercial role is no shield, that supplier complicity is provable in principle — and that liability remains fragmented, national and slow.

Method

Trial Trial: putting the framework under pressure

Trial Trial finding
A research finding from the simulated trial — a structured exercise, not a court judgment.

In 2026, n-Space staged Trial Trial at Somerset House Studios: a public-interest simulated criminal trial of a fictionalised case assembled from documented reporting about AI-assisted targeting. Practising lawyers argued it; expert witnesses were cross-examined; a citizen jury deliberated.

A simulation cannot convict anyone, and this one did not try. Its purpose was adversarial stress-testing: when a legal framework must survive cross-examination — with a named defendant, specific evidence and a jury asking plain questions — its gaps stop being abstract.

The exercise asked when a supplier's contribution becomes legally relevant; what happens when the state controls the strike but the supplier controls parts of the decision environment; what credible notice should oblige; whether the system's role could be reconstructed after civilian deaths; and whether nominal human approval survives realistic tempo. The findings below are research findings, not judgments.

Trial Trial was a public-interest simulated criminal trial conducted by n-Space at Somerset House Studios in 2026. It used a fictionalised case assembled from documented reporting about AI-assisted targeting, and brought practising lawyers, expert witnesses and a citizen jury into a structured adversarial process.

The method matters more than the verdict. Policy discussion tolerates vagueness; cross-examination does not. When a barrister asks a witness exactly which version of the software was deployed, who saw the uncertainty display, and what the supplier did after notice — and the framework has no answer — the gap is exposed in a form a jury, and a committee, can see.

The courtroom as a test environmentClaimEvidenceCross-examinationResponsibilityRemedyEach step forces precision the policy debate can avoid.
The courtroom as a test environment.

The exercise asked

  • When does a supplier's contribution become legally relevant rather than incidental?
  • What happens when the state controls the strike but the supplier controls parts of the decision environment?
  • What should credible notice of harmful or out-of-envelope use require a supplier to do?
  • Can the current framework reconstruct the system's role after civilians are killed?
  • Does nominal human approval survive realistic operational tempo?

Findings

Trial Trial finding
A research finding from the simulated trial — a structured exercise, not a court judgment.
  • State duties remain central; supplier duties supplement rather than dilute them.
  • Supplier responsibility is fragmented across contract, regulation and criminal law, and difficult to establish in any of them.
  • Internal due diligence is usually not independently testable — the court had to take policies on trust or discard them.
  • Human-control claims are often asserted rather than measured; under cross-examination, assertion collapsed quickly.
  • Missing logs, changing software and secrecy can defeat later reconstruction, whatever the applicable law.
  • Adversarial simulation can reveal governance failures before deployment — cheaply, publicly and without waiting for casualties.

The case tried in Trial Trial was fictionalised. Nothing in the exercise establishes wrongdoing by any real company, person or state. Findings are research observations about the legal framework, not verdicts.

SOURCES [20] Trial Trial — original project documentation, n-Space, Somerset House Studios

Trial Trial in full →

Four patterns

Four legal stress tests

Hypothetical
A stipulated scenario used to test the framework. It asserts no fact about any real person or company.

Four recurring factual patterns do most of the damage to the current framework. Each is presented the same way: the pattern, the law engaged, the evidence that goes missing, what currently happens, and the safeguard that would change the answer.

They are deliberately mundane. None involves a rogue robot; all involve ordinary institutional behaviour — throughput pressure, software updates, commercial awkwardness, incomplete records — meeting rules written for different objects.

1 · The human remains in the loop, but the loop is too fast

The pattern
A targeting system produces more recommendations than a person can realistically investigate. Review time falls to seconds; approval rates climb; the record shows a human decision for every strike.
Framework engaged
IHL precautions in attack; UK policy on context-appropriate human involvement. Both assume the human decision is real.
What goes missing
Review time per decision; what was displayed, including uncertainty; refusal rates; performance under realistic tempo — none routinely measured or preserved.
Current outcome
The formal record satisfies the policy; nobody can say whether the control it describes existed. After harm, the operator's keystroke is the best-documented fact in the chain.
Safeguard
Proposed reform — not current law
A policy recommendation. It describes what the editors argue the law should require, not what it requires now.
Make human control measurable: test review time, information quality, uncertainty presentation and authority to refuse at expected operational tempo, as an acceptance condition and continuously in service.

SOURCES [2] Ambitious, Safe, Responsible: our approach to the delivery of AI-enabled capability in Defence, UK Ministry of Defence[15] Technological Evolution on the Battlefield, Center for Strategic and International Studies (CSIS)[20] Trial Trial — original project documentation, n-Space, Somerset House Studios

2 · The reviewed system changes

The pattern
A model, data pipeline, decision threshold or autonomy function changes after procurement and legal review. The product name stays the same; the behaviour does not.
Framework engaged
Article 36 review; procurement acceptance; export licences that assessed an earlier configuration. UK policy accepts iterative review may be needed.
What goes missing
A definition of 'material change'; version identifiers tied to deployments; regression tests against the reviewed baseline; any record that re-review was considered.
Current outcome
The state operates a system whose approval describes something it no longer is. If challenged, neither the reviewed baseline nor the deployed version may be reconstructable.
Safeguard
Proposed reform — not current law
A policy recommendation. It describes what the editors argue the law should require, not what it requires now.
Defined materiality thresholds that trigger re-testing and renewed legal review; deployed-version identification as a contractual and review requirement.

SOURCES [1] Additional Protocol I to the Geneva Conventions, Article 36 (new weapons), ICRC treaty database[4] Government response to the House of Lords AI in Weapon Systems Committee report, UK Government[13] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

3 · The supplier is put on notice

The pattern
A supplier receives credible information that its system is being used outside the agreed operating envelope, or in operations involving possible serious violations. Support and updates continue.
Framework engaged
Contract (if an envelope was defined); export licensing (for continuing transfers); human-rights due diligence (soft law); criminal complicity only at a high threshold of knowledge and contribution.
What goes missing
What notice was received and when; the deployed version and configuration; what technical and contractual control the supplier retained; what it did next.
Current outcome
No clear rule says what notice obliges. Companies improvise between commercial pressure and legal uncertainty; whatever they choose is untestable from outside.
Safeguard
Proposed reform — not current law
A policy recommendation. It describes what the editors argue the law should require, not what it requires now.
A notice-to-action ladder — preserve and assess → report and restrict → escalate or suspend → investigate and remedy — as a continuing statutory and contractual duty, with safeguards against making suppliers private foreign-policy authorities. Proposed reform except where a specific step is already required by an existing contract or licence condition.

SOURCES [12] UN Guiding Principles on Business and Human Rights, UN Office of the High Commissioner for Human Rights[8] Bribery Act 2010, section 7 — failure of commercial organisations to prevent bribery, UK Parliament (legislation.gov.uk)[20] Trial Trial — original project documentation, n-Space, Somerset House Studios

4 · Civilians are killed and the evidence does not survive

The pattern
After a strike kills civilians, logs are incomplete, updates unrecorded, knowledge distributed across institutions, and classified or commercially confidential material blocks public reconstruction.
Framework engaged
The IHL duty to investigate; domestic criminal law; inquests and judicial review; state responsibility and reparation. All depend on evidence none of them created.
What goes missing
Everything at once: recommendation logs, version identifiers, interface records, review times, the chain from data to decision.
Current outcome
Investigations close without findings; families receive conclusions rather than explanations; each institution accurately describes the limits of what it could establish.
Safeguard
Proposed reform — not current law
A policy recommendation. It describes what the editors argue the law should require, not what it requires now.
Tamper-evident logging as an acceptance condition; preservation duties with adverse-inference consequences for unexplained absence, subject to fair-trial safeguards; disclosure calibrated to security rather than defaulting to silence.

SOURCES [13] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)[20] Trial Trial — original project documentation, n-Space, Somerset House Studios[3] Proceed with Caution: Artificial Intelligence in Weapon Systems (HL Paper 16), House of Lords AI in Weapon Systems Committee

The responsibility map

Who may answer after harm?

Current law
A statement of law presently in force, as the editors understand it.

Nine kinds of actor sit around a single strike, and several forms of responsibility can attach to the same event without cancelling one another: state responsibility under international law, command responsibility, individual criminal liability, corporate liability, contractual and regulatory consequences.

The map below shows, for each actor, the core responsibility, the knowledge-or-control question an investigator would ask, the evidence that answers it, and the form accountability could take.

Two errors to avoid. Responsibility is not a single parcel passed down the chain until it reaches whoever touched the interface last. And ordinary engineers and employees are not liable merely because they worked on a product — individual liability follows personal knowledge, conduct and authority.

Concepts

Responsibilities that coexist

State, command, individual, corporate and contractual responsibility can all attach to one event; none cancels another.

The state's duty to conduct lawful operations is non-delegable. A commander's responsibility for precautions does not absorb a supplier's responsibility for honest disclosure, and neither erases an individual's liability if they knowingly assisted a crime. Investigations go wrong when one finding is treated as closing the others.

SOURCES [6] Rome Statute of the International Criminal Court, Articles 25 and 28, International Criminal Court[7] International Criminal Court Act 2001, UK Parliament (legislation.gov.uk)

  • The state

    Lawful use of force; IHL compliance; investigation; reparation.

    Did the state ensure the systems it relies on can be used, reviewed and reconstructed lawfully?

    Evidence · Weapons reviews, procurement conditions, operational records, investigation files.

    Accountability · State responsibility under international law; political accountability; reparation.

  • Political & procurement authorities

    Authorisation, procurement conditions, resourcing of assurance, oversight.

    Was a foreseeable risk accepted without requiring safeguards that were available?

    Evidence · Business cases, contract terms, ministerial submissions, warnings received.

    Accountability · Ministerial and parliamentary accountability; maladministration findings; in extremis misconduct offences.

  • Commanders

    Mission design, target approval, supervision, rules of engagement.

    Were precautions and control adequate in context — and did the commander act on what they knew or should have known?

    Evidence · Orders, ROE, the decision picture as presented, reporting up and down.

    Accountability · Command responsibility; court martial; administrative action.

  • Operators

    Use within training and approved limits.

    Did the operator understand the system's limits — and did the system give them a real decision to make?

    Evidence · Interface records, review time, training logs, what was displayed.

    Accountability · Individual criminal or disciplinary liability only for genuine personal fault — not by default.

  • The supplier corporation

    Design claims, testing, disclosure of limitations, updates, continuing support.

    What did the company know, contribute and have the power to prevent?

    Evidence · Design records, test results, disclosed limitations, update history, notice received and actions taken.

    Accountability · Contractual remedies; regulatory consequences; civil claims; corporate criminal liability where national law provides it.

  • Executives & senior managers

    Decisions within their actual authority: what to sell, support, disclose, continue.

    Did they authorise, conceal, or continue conduct despite known and substantial risk?

    Evidence · Board and management records, briefings, risk assessments, escalations received.

    Accountability · Individual criminal liability on proof of the offence elements; disqualification; employment consequences.

  • Engineers & other employees

    Duties proportionate to role, knowledge and control; raising known risks through protected routes.

    Did a specific person knowingly make a substantial contribution to unlawful use — or report the risk?

    Evidence · Personal role, actual knowledge, what they raised and to whom.

    Accountability · Liability only for personal knowledge and conduct. Working on a product is not an offence.

  • A national assurance authority (proposed)

    Independent testing, certification, change control, incident review.

    Were claims verified, changes controlled, and warnings acted on?

    Evidence · Certification records, audit findings, re-review triggers and outcomes.

    Accountability · Public-law accountability for the regulator itself; judicial review; parliamentary oversight.

  • An international verification body (proposed)

    Common standards, confidential registration, inspection, incident investigation.

    Can states' and suppliers' compliance claims be independently checked?

    Evidence · Declarations, inspection reports, incident investigations.

    Accountability · Findings, certification consequences, referral to competent authorities.

Several forms of responsibility can coexist on the same event. Ordinary employees are not liable merely because they worked on a product.

Political economy

The accountability gap in procurement

Reported practice
Attributed reporting that has not been fully or independently verified.

The hard part is not drafting duties; it is the structure of the relationship. Governments depend on strategic suppliers for capability and technical knowledge. Procurement, national-security and industrial policy all reward preserving those relationships. Classification and commercial confidentiality limit what outsiders can inspect. Rapid procurement can lock a supplier in before assurance capacity matures — and when harm occurs, liability drifts toward the operator, the actor with the least upstream control.

The counterargument deserves stating fairly: duplicative, poorly designed regulation delays urgent capability, favours incumbents who can afford compliance, and excludes smaller firms. Maximal paperwork is not the goal.

The answer this site argues for is lean, technically literate assurance tied to risk and function — designed so that clear rules reduce uncertainty for government and responsible industry alike, and make lawful innovation easier to demonstrate, not harder.

Proposed reform — not current law

A proposed UK accountability architecture

Proposed reform — not current law
A policy recommendation. It describes what the editors argue the law should require, not what it requires now.

Eight practical measures, built on the levers government already holds. Procurement is the first enforcement layer; criminal law is the backstop, not the plan.

Together they make one change: they convert assertions — about human control, testing, updates and supplier conduct — into evidence that exists before anyone needs it.

PROPOSED REFORM — NOT CURRENT LAW. These measures describe what the editors argue procurement and statute should require.

1 · Auditable procurement

High-risk contracts require the evidence base up front: design documentation, data provenance, model and version records, operating-envelope definitions and access for independent testing.

  • What is not contracted for will not exist when it is needed.
  • Documentation scales with risk and function — not maximal paperwork for everything.
2 · Measurable human control

Test review time, information quality, uncertainty presentation, authority to refuse and performance at expected operational tempo — as acceptance conditions and in service.

  • A human in the loop is a claim; these are its measurements.
3 · Continuous legal review

Article 36 review plus defined thresholds for material software change, with re-testing and recertification when they are crossed, and deployed-version identification throughout.

4 · Continuing supplier duty

When credible evidence of harmful out-of-envelope use emerges: preserve and assess, report and restrict, escalate or suspend through defined legal process, cooperate with investigation and remedy.

  • Safeguarded so suppliers do not become private foreign-policy authorities: pre-agreed triggers, expedited regulator or court process, an emergency route for imminent grave harm.
5 · Responsibility allocation

Contracts identify, before deployment, which actor controls data, updates, deployment settings, incident reporting and remedial action — so investigation allocates rather than excavates.

6 · Independent assurance

A technically competent assurance function that neither the procuring team nor the supplier controls, with access, clearances and the authority to require re-review.

7 · Evidence and remedy

Tamper-evident logs, protected records, whistleblower safeguards, incident disclosure calibrated to security needs, and meaningful routes to investigation and victim remedy.

8 · Graduated enforcement

Remediation, contract suspension, fines, debarment, senior-manager consequences — and criminal investigation where the relevant evidential and fault thresholds are met.

  • Procurement is the first enforcement layer. Criminal law is a backstop, not the only accountability mechanism.

Proposed reform — not current law

National assurance and international verification

Proposed reform — not current law
A policy recommendation. It describes what the editors argue the law should require, not what it requires now.

National assurance needs an institution: a technically competent function, controlled by neither the procuring team nor the supplier, with access to design records, versions and test evidence, and authority to require re-review when material changes occur.

Beyond it, this site outlines an International Autonomous Weapons Agency — an IAEA for high-risk military-AI functions, but not a copy of it. Its remit would be narrow: target selection and prioritisation, engagement without a new human decision, lethal swarms, machine-speed coupling, and material changes to certified systems. Its powers would be bounded: standards, confidential registration, review of national processes, laboratory accreditation, inspections, incident investigation, export verification, sanitised findings, referral.

Its limits must be stated as plainly as its promise: non-parties, secrecy, proxy use, covert software, unequal capacity — and the standing risk that certification becomes political legitimation. It does not exist, and nothing here presents it as inevitable.

Concepts

Certification is not immunity

An approval evidences compliance at a moment in time; it cannot excuse concealment, out-of-envelope use or unreviewed material change.

Certification has a scope — a version, an envelope, a set of disclosed facts. Treating a passed test as a shield is the standing failure mode of assurance regimes, national or international. Treated as a floor, with continuing duties attached, it keeps the paper connected to the deployed machine.

SOURCES [13] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

PROPOSED REFORM — this agency does not exist and is not presented as inevitable.

An IAEA for high-risk military AI functions — but not a copy of it

The nuclear analogy earns its place for four things: independent declarations, inspection practice, common technical methods and confidence-building. Its limits are equally instructive: software can be copied, updated and hidden; military AI is dual-use; there is no scarce material to count; non-state actors matter; and lawfulness usually turns on context of use, not possession.

The IAEA asks whether declared nuclear material and facilities match reality. This agency would ask what decisions a system is permitted to make, in which conditions, which version was deployed, and whether a harmful use can be reconstructed.

The remit would be limited to defined high-risk functions:

  • target selection and prioritisation;
  • engagement without a new human decision;
  • lethal swarm coordination;
  • machine-speed coupling of warning, target generation and response;
  • material changes to certified systems.
Bounded possible powers:
  1. minimum technical standards for testing, logging, change control and incident reporting;
  2. confidential registration of systems, envelopes and deployed versions;
  3. review of national assurance and weapons-review processes against the treaty minimum — not their replacement;
  4. accreditation and audit of independent test laboratories;
  5. agreed routine inspections and tightly governed challenge inspections;
  6. incident investigation with protected access to logs and version histories;
  7. export and end-use verification support;
  8. sanitised public findings;
  9. referral of evidence to competent national and international authorities.
Limits, stated plainly
  • non-parties and proxies remain outside the regime;
  • secrecy and covert software limit verification;
  • unequal capacity risks making assurance a rich-state monopoly unless assistance is funded;
  • certification can drift into political legitimation of approved systems — the regime's own standing hazard.

SOURCES [21] IAEA safeguards and verification, International Atomic Energy Agency[5] ICRC position on autonomous weapon systems, International Committee of the Red Cross

For government and Parliament

Recommendations

Proposed reform — not current law
A policy recommendation. It describes what the editors argue the law should require, not what it requires now.

The committee-facing recommendations are listed in full on the recommendations page; the eight measures above are their substance. What they amount to is a single test that procurement, ministers and Parliament can each apply.

“Fit for purpose” must include “capable of lawful use.” The Government–industry partnership is advancing faster than the architecture needed to test, trace and govern it. The answer is not to stop innovation, or to outsource legal judgment to suppliers. It is to build accountability into procurement: measurable human control, review of material software changes, auditable evidence, continuing supplier duties and scrutiny independent of both supplier and procurer. These safeguards protect civilians, the state, responsible companies and the legitimacy of UK defence policy.

Legal content on this site is editorial. Reforms marked PROPOSED REFORM are recommendations, not current law. All legal content requires review by qualified UK public-law, IHL, export-control and international-criminal-law counsel before publication.