The Law Has Unit Tests
A test of a statute looks like a test of code, with three differences: silence is a passing result, a test can expect a question instead of an answer, and every test carries a date.

Here is a test. If you have ever written one for software, you have read this file before:
// Law 12.2, 2024/25 edition: the goalkeeper held the ball for nine seconds.
// Date of law 1 September 2024: the 2024/25 book applies.
test "ifab-01-gk-nine-seconds-2024-25" {
given {
context { legal_time @2024-09-01; timezone "UTC"; }
assert "a0": goalkeeper_within_own_penalty_area(match, i1) { origin case_input; }
assert "a1": goalkeeper_control_seconds(match, i1, 9.0) { origin case_input; }
}
evaluate truth(indirect_free_kick_offence(match, i1));
expect truth_status == TRUE_ONLY;
}Given, when, then. Two facts from the match report, one question to the Laws of the Game, one expected answer. The file sits in the tests/ folder of the package that models the IFAB Laws of the Game, next to the rules it exercises. Run it, and Arxo executes the model of Law 12 against these facts and compares the answer with the expectation.
This is not a metaphor. On 21 September 2026 the corpus of executable canons holds 21,948 such scenario files: contract law, civil codes, insurance statutes, the UN Charter, chess and football, the CSS cascade. They are written in one language, they run under one engine, and they fail the same way a broken unit test fails.
The interesting part is where a test of a law stops behaving like a test of code.
1. Silence is a passing result
In software, a function that returns nothing has a bug. In law, a norm that says nothing about your facts is the correct outcome, and a test has to be able to expect it.
The Declaration by United Nations of 1 January 1942 binds each signatory to employ its full resources against the Axis members "with which such government is at war". The Soviet Union signed. It was at war with Germany and not at war with Japan. Does the Declaration oblige the USSR against Japan?
test "un-test-pair-not-at-war" {
given {
context { legal_time @1942-01-01; timezone "Europe/London"; }
assert "un-war-de2": at_war(UnionOfSovietSocialistRepublics, Germany) { origin case_input; }
}
evaluate truth(employs_full_resources_against(UnionOfSovietSocialistRepublics, Japan));
expect truth_status == NEITHER;
expect evaluation_status == COMPUTED;
}The expected answer is NEITHER: the proposition is neither established nor refuted. The computation finished (COMPUTED), nothing went wrong, and the Declaration is simply silent for this pair. The model does not turn missing facts into "false", because the world of a case is open: nobody asserted that the USSR was at war with Japan, and nobody asserted that it was not.
The same expectation appears in commercial law. Under the UNIDROIT Principles, a non-performance is fundamental if it was intentional, unless the non-performing party would suffer disproportionate loss. A scenario with an intentional breach and a disproportionate loss expects NEITHER for "fundamental non-performance": the intention does not carry the day, and no other ground was pleaded.
A test framework that only knows true and false cannot express this. The test language for canons has to, or half of the law becomes untestable.
2. A test can expect a question
Some elements of a norm are not facts. "Deliberately played the ball", "disproportionate loss", "obvious goal-scoring opportunity": the text hands these to a person with authority to decide. A model that guessed them would be wrong by design. So the model asks, and a test can expect the asking.
Law 11 of football: a player in an offside position receives the ball from an opponent. Whether the opponent deliberately played it, or the ball merely rebounded, is the referee's judgement. Without that judgement the question "offside offence?" has no answer yet, and the test says so:
test "ifab-07-offside-rebound-requires-judgment" {
given {
context { legal_time @2026-09-01; timezone "UTC"; }
assert "a0": part_in_opponents_half(match, i3, p1) { origin case_input; }
assert "a2": part_nearer_goal_line_than_second_last_opponent(match, i3, p1) { origin case_input; }
assert "a4": ball_came_from_opponent(match, i3) { origin case_input; }
assert "a5": played_or_touched_the_ball(match, i3, p1) { origin case_input; }
}
evaluate truth(offside_offence(match, i3, p1));
expect evaluation_status == REQUIRES_JUDGMENT;
expect judgment_request(opponent_deliberately_played_the_ball, MatchReferee);
}The expectation is precise: the evaluation stops with the status "requires judgement", and the open question is this predicate, addressed to this organ, the match referee. Not "the model was unsure". The next scenario in the same folder adds the referee's answer, marked origin adjudicated rather than origin case_input, and expects a definite result with zero open questions.
The same channel runs through the commercial canons. In the UNIDROIT scenario on fundamental non-performance by deprivation, the expected status is TRUE_ONLY for the fundamental character and REQUIRES_JUDGMENT, because the question of disproportionate loss under Article 7.3.1(2)(e) has not been put to the tribunal. Two expectations on one answer, and both must hold.
Civil codes use it too. In the Code civil case below, the harmful act, the damage and the causal link are on file: liability under article 1240 comes back as a question for the judge, because it needs a finding of fault, while prescription on the same facts is established, because it needs only dates.
3. Every test carries a date
The file at the top of this article expects an indirect free kick. Its neighbour holds the same two facts and expects a corner kick:
// The same incident on 1 September 2026: the 2026/27 book applies.
test "ifab-02-gk-nine-seconds-2026-27" {
given {
context { legal_time @2026-09-01; timezone "UTC"; }
assert "a0": goalkeeper_within_own_penalty_area(match, i1) { origin case_input; }
assert "a1": goalkeeper_control_seconds(match, i1, 9.0) { origin case_input; }
}
evaluate truth(restart_awarded(match, i1, corner_kick));
expect truth_status == TRUE_ONLY;
}Nothing in the package was edited between the two tests. Both rules live in the model, each pinned to the edition of the Laws that contains it, each edition with its window of force:
@source(LOTG_2425_LAW_12)
rule GoalkeeperSixSecondsIndirectFreeKick(m, i, secs) strict {
when goalkeeper_within_own_penalty_area(m, i)
and goalkeeper_control_seconds(m, i, secs)
and (secs > 6.0);
then indirect_free_kick_offence(m, i);
}
@source(LOTG_2526_LAW_12, LOTG_2627_LAW_12)
rule GoalkeeperEightSecondsCornerKick(m, i, secs) strict {
when goalkeeper_within_own_penalty_area(m, i)
and goalkeeper_control_seconds(m, i, secs)
and (secs > 8.0);
then restart_awarded(m, i, corner_kick);
}The date of law in the test selects the edition; the edition selects the rule. Of the 21,948 scenarios in the corpus, 19,358 name a date of law explicitly. A test without a date would be a test of "the law in general", and there is no such thing.
Tests are written from the text, not from the model
A unit test that was generated from the implementation tests nothing: it will agree with whatever the code does. The same trap exists for law, and it is easier to fall into, because the model is often the first precise reading of the text anyone has produced.
So the rule for canons is the rule for good test suites: the expectation comes from the source, and it is frozen. A change in the model has to break scenarios; it never rewrites them. When a scenario and the model disagree, the text is opened, the reading is settled, and the losing side is fixed.
An example of what the text can hide. Article 19(4) of the Kazakh compulsory motor insurance law adds a coefficient of 0.8 "for other towns and settlements in the regions listed in paragraph 3". The table in paragraph 3 lists regions, and it also lists Almaty, Astana and Shymkent, which are cities of republican significance, not regions. A scenario written from the text for "Almaty, other settlement" expects the coefficient not to apply. A scenario written from the table would expect the opposite. Three independent readers of the text agreed with the first expectation, and that is the one the package holds.
Continuous integration for a bill
Once a statute has a test suite, the question every engineer asks becomes available to a legislator: what does this change break?
Arxo answers it with an impact report. Take the current model of an act and the model of the draft amendment, run every scenario and every declared goal against both, and compare:
- which norms changed in meaning, counted on the structure of the rules rather than on their wording, so that a reworded label with the same logic is reported as no change;
- which scenarios changed their outcome, with the answer before and after;
- which goals the act declares still hold, and which one held before the amendment and fails after it;
- which of the amended norms are executed by no scenario at all.
That last line is the one a legislator will not get from a reading. The report names it plainly: for those norms, "outcome unchanged" is an absence of observation, not a conclusion. The test suite of a statute has coverage, and the coverage has holes, and the holes are listed.
The same report runs on the axis of time instead of the axis of drafts: two dates of law, one act, and the scenarios show what the passage of an edition changed for the cases people actually brought.
Where to look
The UNIDROIT Principles of International Commercial Contracts are published as executable packages at github.com/arxohq/arxo-unidroit-picc. Each package carries its tests/ folder; the scenarios on fundamental non-performance quoted above are in the termination file of the non-performance package. Open one, read the facts, and if you believe the Principles answer differently, send the counterexample as an issue in the same repository. That is how a scenario enters the suite.
Arxo executes the canon. Models formalize; Arxo executes and proves. The tests are how you know which is which.