SatQuery / docs /TESTING.md
thundercode's picture
release: add docs/TESTING.md
55f002c verified
|
Raw
History Blame Contribute Delete
96.9 kB

SatQuery AI — Testing, Verification and Evidence Integrity

Status of this document. This is the testing chapter of the public SatQuery AI release. It enumerates the actual suites that exist, states what each one pins, and reports the full-suite result including its failures. Every test file, count, command and code excerpt below was read from the repository. Where a value could not be established it is written UNKNOWN — not established from the available evidence.

Status vocabulary (see DOCS_STYLE_GUIDE.md §2): IMPLEMENTED · VERIFIED · MEASURED · ATTEMPTED · NOT RUN · BLOCKED · DEFERRED · REJECTED · OPEN · RESOLVED · CLOSED.

The one rule this chapter follows without exception: a green suite is not evidence about a property nobody wrote an assertion for, and a failure is never reclassified as a pass. The project learned this the hard way twice — a guard that pinned a defect as the specification (docs/DEPLOYMENT_ARCHITECTURE.md §5.5) and a guard whose assertion was satisfied by an upstream fix so it could not see the downstream site (F-15b). Both are recorded below rather than smoothed over.


Table of contents


0. How to read this document

0.1 The three evidence classes

The repository encodes an evidence class on every test via a pytest marker (pytest.ini):

markers =
    unit: fast, no I/O, no server, no network. Evidence about a component in isolation.
    integration: exercises two or more real components wired together in-process.
    smoke: a minimal end-to-end path run against real local artifacts, not a mock.
    real_inference: ran the real model on real inputs in this environment.

The comment above the markers states the purpose exactly, and it is the reason the markers exist rather than being decorative:

"The markers exist to make a test's EVIDENCE CLASS explicit: a unit test never claims a deployment was exercised, and no test may be reported as 'real inference' unless it ran the real model on real data in this environment."

And one deliberate non-marker, quoted in full because it is a subtle but important choice:

"environment_blocked is not a test marker: a blocked path is reported in the STEP 7 report rather than encoded as a permanently-skipped test, because a skip can be mistaken for coverage."

That sentence is the spine of this chapter. A skip that looks like coverage is the failure mode the marker design is built to avoid.

0.2 What "verified" means here

Term Means Example
Unit one component, no I/O, no server, no network tests/unit/test_evidence_engine.py
Integration two or more real components wired in-process tests/integration/test_step7_backend_chain.py
Smoke a minimal end-to-end path over real local artifacts tests/unit/phase6_smoke.py (standalone)
Real inference the real model on real inputs, in this environment — see §11
Live validation the deployed stack driven by a browser §8

The distinction matters most at the boundary between integration and live. docs/API_CONTRACT.md §8 states the boundary plainly: "no test in this repository dials a network address", and it points at docs/ITEM5_INTEGRATION_SUITE_SCOPE.md, which "records what the 26-test integration suite proves (the app's boundary, in-process) and what only a live deployment can prove (reachability, cold start, memory ceilings)." A green tests/integration run is therefore not evidence about a deployment.

0.3 The two defects that shaped the discipline

Two incidents are referenced repeatedly below, because they are why this chapter is written the way it is:

  1. A guard that pinned the defect. The test tests/unit/test_controller.py::test_trace_records_inputs_and_query asserted envelope.trace.inputs == list(request.assets) — which, after the upload handler rewrote handles into paths, pinned a filesystem-path disclosure as the specification. The lesson, recorded in docs/DEPLOYMENT_ARCHITECTURE.md §5.3: "A guard is only as strong as the behaviour it pins, and this one pinned the defect — a reminder that a green suite is not evidence about a property nobody wrote an assertion for."
  2. A guard that could not fail. The to_trace() seam's scrub had no covering test: falsifying it alone left the producer-side guard GREEN, because the producer had already scrubbed the value the guard inspected. Recorded as F-15b in docs/DEPLOYMENT_ARCHITECTURE.md §5.5, and remedied by test_to_trace_scrubs_a_detail_that_arrives_from_any_constructor, which constructs the entry directly through the constructor.

Both are the same shape: a guard whose assertion is satisfied by something other than the code under test.


1. Testing philosophy: tests pin contracts, not implementation

1.1 The philosophy, as the suites state it

The repository's test docstrings state a consistent philosophy, and the consistency is the point. Four representative statements, each quoted from a module docstring:

Doc-guards pin examples, not prose (tests/unit/test_api_contract_doc.py):

"These tests validate the EXAMPLES, not the prose. They cannot detect a wrong sentence (e.g. a mis-stated latency budget). What they guarantee is that a frontend developer who copies an example verbatim gets a request the server accepts -- which is the failure mode that actually costs time."

Guards pin the route, not the validity (tests/unit/test_frontend_live_wiring.py):

"It also pins the chang word-boundary defect: \bchang\b cannot match 'changed', so the page's own default question ('What changed here?') routed to vqa instead of the change path. That is invisible to any test that only checks 'is this a valid enum value' — both branches produce a valid value. The test therefore asserts the ROUTE, not just the validity."

Tests read shipped source when there is no importable symbol (tests/unit/test_frontend_live_wiring.py):

"Why these are read from the source text: the frontend is plain ES5 with no build step and no module exports, so there is no importable symbol. Parsing the source is the only way to assert the shipped bytes. The parsing is anchored on named tokens rather than line numbers, so it does not silently pass if the file is restructured."

A runbook must not assert facts the repository contradicts (tests/unit/test_runbook_doc.py):

"A runbook is worse than useless if a copied command fails or a quoted number was never measured: the operator concludes the system is broken rather than the document."

1.2 The three anti-patterns the suites are written to avoid

Anti-pattern What it looks like How the suites avoid it
Pinning the defect an assertion that encodes today's buggy behaviour as correct guards are falsified before being trusted: the fix is reverted and the guard is shown to fail (§1.3)
Vacuous guards an assertion satisfied by an upstream layer, so the downstream site is untested the seam is tested by constructing the object directly through its constructor (F-15b)
Measurement by prose a text search that is tripped by documentation about a defect guards walk the parsed examples, not the document text (test_api_contract_doc.py::test_no_contract_example_fabricates_an_artifact_uri)

The third is worth quoting in full, because it is a rule about what a guard must fail on (tests/unit/test_api_contract_doc.py):

"This walks the PARSED examples rather than the document text on purpose. The prose in section 2.4 deliberately quotes the old artifact://run/9f2c.../change_map.png example when explaining why it was removed, so a text search would be tripped by documentation about the removal -- which is measuring the wrong thing. A guard must fail on the defect, not on a description of the defect."

1.3 Falsification is part of the discipline

A guard is not trusted until it has been shown to fail against the defect it guards. The pattern is recorded with measurements throughout the project. The clearest instance is the F-15 four-site falsification (docs/DEPLOYMENT_ARCHITECTURE.md §5.5), where each site was reverted individually and the guards re-run:

Site reverted test_a_construction_failure_publishes_no_filesystem_path test_to_trace_scrubs_a_detail_that_arrives_from_any_constructor
registry.py:597 (_failure_entry) FAILS passes (correctly — it tests the seam, not the producer)
registry.py:309 (to_trace) passes — blind FAILS

The "passes — blind" cell is the F-15b finding: the producer guard cannot see the seam, because the producer has already scrubbed the value. The two guards are therefore not duplicates, and the table is the evidence.

A second falsification table covers the two controller sites (docs/DEPLOYMENT_ARCHITECTURE.md §5.5):

Site reverted test_a_failed_step_publishes_no_filesystem_path — assertion that fires
controller.py:801 (_failure_evidence) line 590, payload["message"]
controller.py:969 (_warnings) line 582, warnings[]

And a third, for the trace-path fix (docs/DEPLOYMENT_ARCHITECTURE.md §5.4):

Step Result
Revert only the PARSE site, run the module 1 failed, 55 passed at tests/unit/test_controller.py:1153
Restore byte-exact, verify hash 0c00ae6f54300668122728dc706344044bc19d48328ceead30323c7147e04397
Module re-run 56 passed
Is the guard vacuous? No — the shape check only discriminates because the fixture writes real GeoTIFFs, and the name-equality check rejects the revert independently

The "restore byte-exact, verify hash" row is the anti-tamper step: after reverting a site to falsify a guard, the file is restored and its hash checked, so the falsification itself cannot leave the tree in a modified state.

1.4 Determinism is asserted, not hoped for

Several suites pin determinism as a first-class property rather than a nice-to-have:

Suite Determinism property
tests/unit/test_evidence_engine.py the same claims produce a byte-identical, identically-ordered collection even when uuids differ between processes (§4)
tests/unit/test_router_threshold_sweep.py validation-only sweep contracts
tests/unit/test_change_threshold_sweep.py validation-only sweep contracts
tests/unit/test_report_generator.py "pure, deterministic, fixtures only"
tests/unit/test_vlm_evaluate.py the predeclared acceptance rule
tests/unit/test_fusion_seed_variance.py the predeclared seed-variance pass

The determinism claim that matters most is the evidence engine's, because it is what makes the live validation's verdict-recomputation property hold (§9).


2. The suite inventory

2.1 The layout

pytest.ini sets the roots:

[pytest]
testpaths = tests
pythonpath = .
addopts = -q --tb=short
filterwarnings =
    ignore::DeprecationWarning
    ignore::PendingDeprecationWarning
    ignore::rasterio.errors.NotGeoreferencedWarning

tests/ contains eight directories and two top-level modules:

Location Contents Evidence class
tests/unit/ 98 Python files — 94 collected test modules + __init__.py, conftest.py, _qa_probe_tokens.py, phase6_smoke.py unit (mostly)
tests/e2e/ test_demo.py, test_gate4_e2e.py, _demo_probe.py end-to-end
tests/geospatial/ test_transform.py unit
tests/integration/ test_step7_backend_chain.py, test_step7_r02_serving.py integration
tests/leakage/ test_leakage.py unit
tests/model/ test_vlm_contract.py, test_vqa_quality_gate.py unit / integration
tests/routing/ test_router.py unit
tests/ test_config.py, test_schemas.py unit

tests/unit totals 49,673 lines across its 98 Python files. Two files there are deliberately not collected:

File Why it is not collected
tests/unit/_qa_probe_tokens.py "QA scratch probe (NOT collected by pytest -- underscore prefix)."
tests/unit/phase6_smoke.py "Phase 6 real CPU smoke test -- standalone, NOT collected by pytest."
tests/unit/conftest.py "Shared fixtures for the Phase 6 VLM unit tests." — a fixture module, not a test module
tests/unit/__init__.py package marker

2.2 The full file inventory, with what each covers

The descriptions below are the first line of each module's own docstring — the module's statement of what it pins, not a paraphrase.

2.2.1 Gateway, deployment and configuration

File What it covers (module docstring)
test_gateway_app.py "The gateway ASGI layer: its contract, its allowlist, and real HTTP execution."
test_gateway_assets.py "The ephemeral asset store: opaque handles, three obligations, real lifetimes."
test_gateway_policy.py "The gateway's decisions must be correct, cheap, and never reach the Space on a rejection."
test_gateway_responsibilities.py "STEP 7 §C — the eighteen gateway responsibilities, as UNIT tests."
test_space_app.py "The Space entrypoint must be cheap, honest, and hash-preserving."
test_deployment_adapter.py "The deployment capability adapter: registry-authoritative, contract-shaped."
test_deploy_config.py "Tests for the Phase 18 deploy packaging manifest and its validator."
test_deploy_requests.py "Phase 18 §79 DEPLOYMENT: cold-start and sequential-request behaviour."
test_config.py "Phase 1 tests — configuration registry and frozen contract guards."
test_code_revision.py "Tests for core.code_revision — provenance defect section 6b."
test_packaging_script.py "Regression tests for scripts/package_kaggle_code.py."

2.2.2 Controller, planner, registry and schemas

File What it covers
test_controller.py "Tests for core.controller — dispatch, partial failure and assembly (Phase 15)."
test_controller_seam.py "The controller's asset-resolution seam — regression tests for a real defect."
test_planner.py "Tests for core.planner — the deterministic policy planner (Phase 15)."
test_registry.py "Tests for core.registry — the specialist capability registry (Phase 15)."
test_schemas.py "Phase 1 tests — canonical schema contract."
test_errors.py "Tests for core.errors — the taxonomy, and the F-15 path scrubber."
test_evidence_engine.py "Tests for the evidence engine — the aggregator (Phase 13)." (§4)
test_report_generator.py "Unit tests for reports.generator — pure, deterministic, fixtures only."

2.2.3 Router

File What it covers
test_router_threshold_sweep.py "Phase 4/13 — contracts for the validation-only ROUTER threshold sweep."
tests/routing/test_router.py "Phase 4 tests — the intent router."

2.2.4 Change detection and change-VQA (the largest subsystem by file count)

File What it covers
test_change.py "Phase 9 tests — change detection model, dataset, and post-processing."
test_change_levir_real_layout.py "Phase 9 — the LEVIR-CD loader against the REAL dataset layout."
test_change_phase9_decontamination.py "Phase 9 -- pins that stop the change-detection module from being re-contaminated."
test_change_specialist.py "Phase 9 — the change specialist serving contract."
test_change_threshold_sweep.py "Phase 9 — contracts for the validation-only threshold sweep."
test_change_train_script_contract.py "Phase 9 — contracts for scripts/train_change.py."
test_change_vqa_amp.py "R-02 — the mixed-precision training contract."
test_change_vqa_dataset.py "R-02 test areas E-J — the CDVQA dataset layer."
test_change_vqa_eval_cli.py "R-02 — the evaluation CLI's empty-split diagnosis."
test_change_vqa_features.py "R-02 test areas K-L — frozen change features and the feature caches."
test_change_vqa_head.py "R-02 test areas M-Q — batching, the reasoning head, and the metrics."
test_change_vqa_integration.py "R-02 test area R — the change-VQA specialist and its planner/registry wiring."
test_change_vqa_kaggle_notebook.py "R-02 — the Kaggle notebook's discovery contract."
test_change_vqa_prepare_script.py "R-02 — prepare_record.json must be the provenance of the whole directory."
test_change_vqa_promotion.py "STEP 1 — the promoted R-02 change-VQA head at its canonical serving path."
test_change_vqa_smoke.py "R-02 — the local smoke training run, and the artifact it leaves behind."
test_change_vqa_vocab.py "R-02 test areas A-D — the CDVQA answer space and question ontology."
test_cdvqa_adapter.py "Tests for the CDVQA reader, join and example construction."
test_cdvqa_second_overlap.py "Tests for scripts/check_cdvqa_second_overlap.py."

2.2.5 Optical-SAR

File What it covers
test_optical_sar_croma.py "CROMA wrapper — the interface verified against the real model."
test_optical_sar_croma_geometry.py *"Phase 11 geometry guards — the four requirements that were measured but…"
test_optical_sar_fusion_head.py "Optical-SAR fusion head — the frozen concatenation and channel dropout."
test_optical_sar_prompts.py "Optical-SAR prompts — the narration boundary."
test_optical_sar_radiometry.py "CROMA encoder-input radiometry — the DEV-2 ruling, as executable guards."
test_optical_sar_sensor_adapter.py "Optical-SAR sensor adapter — the two hard rules, enforced."
test_optical_sar_specialist.py "Phase 11/12 — the optical-SAR specialist serving contract."
test_app_serving_optical_sar.py "Regression guards for the optical-SAR serving wiring (Phase 14 / Pass 16)."
test_eval_fusion_115.py "Tests for the pre-registered 11.5 metric tool."
test_fusion_extraction.py "Tests for the Phase 12 frozen-feature extraction pipeline (fixtures only)."
test_fusion_seed_variance.py "Tests for the predeclared seed-variance pass (Gate F D-03 / D-05)."
test_fusion_training.py "Tests for the Phase 12 fusion-head training loop (fixtures only)."
test_reben_adapter.py "Tests for the reBEN v2 -> PairedSample adapter (Phase 12)."
test_analyze_reben_labels.py "Tests for the reBEN label analyser."

2.2.6 Grounding

File What it covers
test_grounding_dataset.py "Tests for the Phase 8 grounding dataset."
test_grounding_degenerate_fallback.py "Regression tests for the degenerate fallback-box guard (Phase 8)."
test_grounding_feature_extraction.py "Tests for the frozen-feature cache."
test_grounding_head_wiring.py "Wiring the trained RemoteCLIP grounding head into production (work order §5)."
test_grounding_metrics.py "Tests for the grounding metrics."
test_grounding_nms.py "Unit tests for grounding NMS (specialists/grounding/head.py::nms)."
test_vrsbench_loader.py "Tests for the VRSBench referring-expression loader."
test_vrsbench_lazy_images.py "Tests for the lazy-image mode of the VRSBench loader."

2.2.7 VLM training and evaluation

File What it covers
test_vlm_artifact.py "Tests for the Phase 6 adapter artifact (training.vlm.artifact)."
test_vlm_collate.py "Tests for Phase 6 batch assembly (training.vlm.collate)."
test_vlm_config.py "Tests for the Phase 6 training recipe (training.vlm.config)."
test_vlm_dataset.py "Tests for the Phase 6 BigEarthNet -> SmolVLM instruction corpus."
test_vlm_evaluate.py "Tests for the predeclared acceptance rule (training.vlm.evaluate)."
test_vlm_evaluate_cache.py "Tests for the resumable per-question answer cache in training.vlm.evaluate."
test_vlm_evaluate_decision_split.py "Tests for decide_acceptance(..., decision_split=...)."
test_vlm_formatting.py "Tests for Phase 6 prompt construction and loss masking (training.vlm.formatting)."
test_vlm_lora.py "Tests for Phase 6 LoRA injection (training.vlm.lora)."
tests/model/test_vlm_contract.py "Phase 5 tests — SmolVLM contract, prompts, and the VQA specialist."
tests/model/test_vqa_quality_gate.py "Integration tests: the F5-5 quality gate blocks the VLM before it runs."

2.2.8 Datasets, metrics and evaluation

File What it covers
test_bigearthnet_blocks.py "Tests for the T2 spatial-block scene key (Phase 12)."
test_bigearthnet_pairing.py "Tests for the BigEarthNet optical<->SAR pairing step (Phase 12)."
test_bigearthnet_prep.py "Tests for the BigEarthNet reader and instruction-pair construction."
test_prepare_script.py "Tests for the BigEarthNet preparation CLI."
test_caption_metrics.py "Caption metrics (ruling R-16): BLEU, ROUGE-L, CIDEr, BERTScore."
test_vqa_metrics.py "Tests for the Tier-1 VQA / Change-VQA text metrics (evaluation.metrics.vqa)."
test_eval_normalize.py "Tests for evaluation.normalize — the plan section 63 metric normaliser."
test_evaluation_runner.py "STEP 5 — the evaluation runner."
test_calibration.py "STEP 2 — the calibration fitter, the artifact, and the consumer wiring."
test_quality_gate.py "Tests for the deterministic input-quality gate (finding F5-5)."
test_benchmark_adapters.py "STEP 5 — benchmark adapters: contract, registry, and the non-invention guard."
test_benchmark_adapters_real.py "The four real-corpus benchmark adapters: correctness against synthetic fixtures."
test_public_test_corpus.py "STEP 5 — the immutable public-test corpus (plan section 37, Test Set Firewall)."
test_manifest_freeze.py "Tests for the frozen dataset-manifest record (evaluation/manifest_freeze.json)."
test_prompt_freeze.py "Tests for the frozen prompt-set record (evaluation/prompt_freeze.json)."
test_run_manifest_population.py "Tests for STEP 4 — runtime population of the run manifest."
test_provenance_fixes.py "Tests for the R-02 provenance fixes: sections 6a, 6b and 6c."
test_diagnose_feature_cache.py "Tests for scripts/diagnose_feature_cache.py's exit-code contract."

2.2.9 Serving composition and the app boundary

File What it covers
test_app_serving.py "Regression tests for app.serving — the public serving composition root."
test_phase6_closure.py "Tests for the Phase 6 closure record."

2.2.10 Contract-conformance, frontend and documentation guards

File What it covers
test_api_contract_doc.py "The API contract must not drift from the schemas it documents."
test_frontend_guide_doc.py "The contract documents must stay consistent with each other and the schemas."
test_frontend_live_wiring.py "Regression tests for the frontend <-> orchestrator live wiring." (§5)
test_runbook_doc.py "The deployment runbook must not assert facts the repository contradicts."
test_deploy_config.py "Tests for the Phase 18 deploy packaging manifest and its validator."
test_step7_api_contract_conformance.py "STEP 7 section D -- the API contract, tested against the schema that implements it."
test_safe_delete_shim.py "The Windows safe-delete shim: verbatim prefixes, and the shell's return code." (§7)

2.2.11 The other directories

File What it covers
tests/geospatial/test_transform.py "Phase 2 tests — geospatial transforms and the raster input contract."
tests/leakage/test_leakage.py "Phase 3 tests — Gate 1: dataset manifests and scene-level leakage isolation."
tests/integration/test_step7_backend_chain.py "STEP 7 sections H, I, M -- the real inference service, driven over HTTP."
tests/integration/test_step7_r02_serving.py "STEP 7 section I -- the R-02 head, verified through the real serving path."
tests/e2e/test_demo.py "End-to-end tests for the Phase-19 demonstration driver (demo.run_demo)."
tests/e2e/test_gate4_e2e.py "Gate-4 end-to-end test — the full control tier in its real local state."
tests/e2e/_demo_probe.py "Subprocess probe: run the demo in a FRESH interpreter and report which…" (helper)

2.3 What the inventory says about coverage shape

Three observations a reader should draw from the inventory:

  1. The largest single cluster is change detection and change-VQA (19 files), which matches the project's stated engineering priority ordering (docs/MASTER_ARCHITECTURE_PLAN.md §2.1: optical-SAR first, change second).
  2. Every capability has a serving contract test, separate from its model tests. For example test_change.py covers the model/dataset/post-processing while test_change_specialist.py covers "the change specialist serving contract"; the same split exists for optical-SAR (test_optical_sar_fusion_head.py vs test_optical_sar_specialist.py) and grounding. The split is deliberate: a model that works in a notebook and a specialist that serves the contract are different claims.
  3. Documentation is tested as code. Five of the 94 modules are doc-guards (§3), and they are counted in the 183-passed doc/frontend suite (§6) rather than being an afterthought.

2.4 The integration suite's documented scope

docs/ITEM5_INTEGRATION_SUITE_SCOPE.md is the record of what the integration suite proves. The contract summarises it (docs/API_CONTRACT.md §8):

"docs/ITEM5_INTEGRATION_SUITE_SCOPE.md records what the 26-test integration suite proves (the app's boundary, in-process) and what only a live deployment can prove (reachability, cold start, memory ceilings). Read it before treating a green tests/integration run as evidence about a deployment — no test in this repository dials a network address, including this contract's own /v1/* examples."

The exact number of tests collected in a single tests/integration invocation is UNKNOWN — not established from the available evidence. The documented figure is 26; the evidence this chapter can establish is the two module files and their stated scope.


3. The doc-guard tests: keeping documentation and code in sync

3.1 Why documentation is tested

Five modules treat documentation as a tested artefact. The rationale is stated in each module's docstring, and the common thread is that a documentation defect costs a real debugging session:

Guard The failure it prevents
test_api_contract_doc.py a frontend builds against an example that does not validate, and the cause looks like a frontend bug
test_frontend_guide_doc.py a value a frontend must send or read is documented wrong, producing a 422 and a wasted session
test_runbook_doc.py an operator copies a command that fails, or quotes a number that was never measured, and concludes the system is broken
test_deploy_config.py the deploy manifest silently moves the frozen Config.hash
test_step7_api_contract_conformance.py the contract and the schema that implements it drift apart

test_api_contract_doc.py records what the guards actually catch, with a list of four real defects the act of writing the contract produced:

"Writing this contract produced exactly that failure four times over -- Box fields were nested instead of flat, Evidence used summary/value/source instead of coordinates/score/source_specialist, CoordinateSystem values were shorthand rather than the real normalized_0_1/pixel/geo, and ExecutionTrace.inputs/outputs were dicts rather than lists. Each was caught by validating the examples against the real models, which is what these tests do permanently."

3.2 test_api_contract_doc.py

Scope: validate every fenced ```json block in docs/API_CONTRACT.md against the real Pydantic models.

Mechanism Detail
Fixture contract_text asserts the document exists, then reads it
Fixture json_blocks regex-extracts every ```json block and parses it; a block that is not valid JSON raises AssertionError naming the block index
_find(blocks, predicate) locates an example by shape rather than by position
Test What it pins
test_every_json_block_in_the_contract_parses "A malformed example is unusable; this fails loudly rather than silently."
test_the_result_envelope_example_validates the envelope example validates against ResultEnvelope; schema_version == SCHEMA_VERSION; confidence.method in ("uncalibrated", "temperature_scaling")
test_no_contract_example_fabricates_an_artifact_uri no parsed example contains an artifact:// URI or a filesystem path in change_map / artifact_ref (walks the parsed examples, §1.2)
(further tests) documented task/coordinate-system values are the real ones; every error code documented; the frozen config hash is recorded; the calibration caveat is recorded

The confidence.method assertion is the honest-signal guard: it pins that the contract's example cannot claim a calibration method the engine does not have. §4 shows the engine side of the same rule.

3.3 test_frontend_guide_doc.py

Scope: docs/API_CONTRACT.md and docs/FRONTEND_INTEGRATION.md must agree with each other and with core/schemas.py.

It imports the real models — Box, CoordinateSystem, Evidence, EvidenceType, Region, Task — and compares the prose against them. The docstring lists the five defects that motivated it:

"Writing them produced, in order: CoordinateSystem documented as normalized/geographic (real: normalized_0_1/geo), Box geometry documented as nested (real: flat), Evidence fields named summary/value/source (real: coordinates/score/source_specialist), ExecutionTrace.inputs/outputs as objects (real: lists), and Task.unsupported missing from both tables."

The representative test asserts a bidirectional requirement:

def test_every_task_is_documented_in_both_documents(api_text, frontend_text):
    for task in Task:
        assert f"`{task.value}`" in api_text, f"Task.{task.value} missing from API contract"
        assert f"`{task.value}`" in frontend_text, (
            f"Task.{task.value} missing from the frontend guide; a value the client "
            f"may receive but cannot find documented is a crash waiting to happen"
        )

The message on the second assertion is the rationale: "a value the client may receive but cannot find documented is a crash waiting to happen."

The suite also pins: no login screen, no secrets in the browser, the confidence-property trap, and serialized analyses (docs/EVALUATION.md §7.2).

3.4 test_runbook_doc.py

Scope: docs/BACKEND_DEPLOYMENT_RUNBOOK.md must not assert facts the repository contradicts.

The module names two failure modes:

  1. "Invented measurements. The runbook quotes command output. If it quotes a byte size, a capability list, a function name or an exit code that the repository does not produce, the quote is fabricated. Where a fact is cheap to verify locally, this module verifies it."
  2. "Invented API surface. The runbook calls into the codebase (build_serving_registry, available(), CHANGE_CHECKPOINT, CHANGE_VQA_HEAD). A runbook that names a function the code does not export is a runbook whose commands raise AttributeError on the operator's first copy-paste."

It reads three documents — the runbook, docs/DEPLOYMENT_ARCHITECTURE.md and docs/API_CONTRACT.md — so that a claim contradicted across documents is caught, not only a claim contradicted by code.

The most important class of assertion in this module is stated in its docstring:

"The absence of a deployment cannot be tested here, so these tests assert the runbook's local claims and, separately, assert that the document itself labels its unimplemented sections as unimplemented. The second class matters more than it looks: a runbook that describes a live deployment without having performed one is the exact overclaim the work order forbids."

So the guard pins honesty about status, not merely factual accuracy: an "undecided" item must be declared undecided, and a "designed" item must not be described as "verified" (docs/EVALUATION.md §7.2).

3.5 test_deploy_config.py

Scope: the Phase 18 deploy manifest and its validator, and the config-hash regression guard.

The docstring states the two properties it proves:

"These tests prove the manifest is INERT — it is never merged into the config registry and cannot move Config.hash — and that the validator rejects each failure mode. Negative cases use in-memory dicts via validate_documents, so the real config files are never mutated."

Key elements:

Element Detail
REGISTRY_MARKER test_deploy_manifest_exists_and_declares_the_non_registry_marker asserts doc[REGISTRY_MARKER] is False
test_validator_reports_no_errors_on_the_real_files assert validate() == []
test_validator_cli_exits_zero_as_a_standalone_script drives the real entry point in a child interpreter (the same one running the suite) "to prove the module's own sys.path bootstrap makes core importable WITHOUT pytest's pythonpath = . shortcut"
GPU_DURATION_KEYS gpu_duration_vqa, gpu_duration_grounding, gpu_duration_change, gpu_duration_optical_sar
FROZEN_CONFIG_HASH imported from tests.test_config rather than re-declared — "We import it rather than re-declare it, so the two cannot drift."

The FROZEN_CONFIG_HASH import is the anti-drift move applied to the guard itself: the deploy-config test does not carry its own copy of 78f1e3700da15aa1, it imports the authority.

The suite also pins the C-8 constraints: torch_compile false, CPU mode required, lazy load, single model cache (docs/EVALUATION.md §7.2).

3.6 test_step7_api_contract_conformance.py

"STEP 7 section D -- the API contract, tested against the schema that implements it." This is the schema-level counterpart to test_api_contract_doc.py: where the doc-guard validates the examples in the markdown, this module tests the contract against the implementing schema directly.

3.7 What the doc-guards cannot do

The modules say so themselves, and the honesty is worth reproducing because it bounds what a green run means:

"Prose can still be wrong in ways these tests cannot see -- a mis-stated latency, a wrong reason for a decision. What is asserted is that every value a frontend must send or read is the real value, which is the class of error that produces a 422 and a wasted debugging session." — tests/unit/test_frontend_guide_doc.py

"These tests validate the EXAMPLES, not the prose. They cannot detect a wrong sentence (e.g. a mis-stated latency budget)." — tests/unit/test_api_contract_doc.py

So a doc-guard guarantees example-level and value-level conformance. It does not guarantee that a sentence is true. That distinction is the reason test_runbook_doc.py exists separately — it targets a different class of claim (numbers, names, and status honesty) that the example-validators cannot reach.


4. The evidence engine: purity, determinism and the reproducibility property

4.1 Why this suite exists at all

The evidence engine is the component that turns a SpecialistResult into a proof: an ordered, identified, digestible EvidenceCollection. Everything the release claims about "evidence" downstream — the trace, the audit trail, the ability to recompute a verdict from a recorded run — rests on one property of this component: the same inputs must produce the same output, byte for byte.

If that property fails, then two runs of the same request can produce two different evidence sets, and no verdict can be recomputed from a record. The evidence engine is therefore the load-bearing component behind §9's "verdict-independent evidence" claim, and its suite is the guard on that load-bearing property.

The suite is tests/unit/test_evidence_engine.py — 73 test functions in 877 lines (tests/unit/test_evidence_engine.py).

4.2 The contract, in the module's own words

The module docstring states exactly two failure modes it is written against, and a third honesty rule:

"The engine's whole value is that it is reproducible and lossless. So the tests are mostly about the two ways that can silently break:

determinism — the same claims must produce the same ordered, identified collection even when specialists finish in a different order, or when the uuids differ between processes. no loss — no specialist's evidence may vanish without being counted, and agreement between specialists must be recorded rather than collapsed with one name thrown away.

Plus the honesty rule on confidence: an uncalibrated engine must pass the raw score through and SAY it is uncalibrated, never manufacture a fitted mapping." — tests/unit/test_evidence_engine.py

Three properties, then: determinism, no loss, and confidence honesty. The suite is organised around them.

4.3 The determinism family

The core purity claim is a single test whose docstring is the whole argument:

def test_aggregate_is_deterministic_across_repeat_calls() -> None:
    """Same input -> byte-identical output. The core purity claim."""

It asserts three things across two calls with the same input — identical ids(), identical evidence_digest(), and identical model_dump() lists. "Byte-identical output" is not a figure of speech: the digest is the serialised form, so a difference anywhere in the collection changes the digest.

The determinism family, read from the file:

Test What it pins
test_aggregate_is_deterministic_across_repeat_calls the core purity claim (line 132)
test_order_is_independent_of_input_order forward vs. backward input order give the same digest
test_equal_scores_order_deterministically_by_coordinates "Ties must not fall back to insertion order — that is input-order dependence wearing a disguise"
test_higher_score_sorts_first_within_a_type the ordering rule inside an evidence type
test_unscored_items_sort_after_scored_items a total order even when some items carry no score
test_types_are_grouped_together type grouping is stable
test_ids_are_sequential_and_zero_padded the identifier format
test_ids_restart_from_one_for_each_aggregation ids are a function of the collection, not global state
test_source_results_are_not_mutated purity — computing evidence does not write back
test_result_objects_are_not_mutated purity, at the result level

The two mutation tests are the purity half of "purity and determinism": the engine is a function, not a transformer of its inputs. A caller's SpecialistResult is the same object after aggregate as before.

4.4 The no-loss family

Losslessness is the harder property to test, because the ways evidence can vanish are subtle. The deduplication tests are where it is pinned:

Test What it pins
test_payload_only_differences_are_deduplicated_as_one_claim two items differing only in payload are one claim
test_identical_items_from_one_specialist_collapse intra-specialist dedup
test_deduplication_ignores_random_uuid dedup is on content, not on the process-local uuid
test_deduplication_ignores_payload_differences_and_merges_them payloads merge rather than one being dropped
test_different_coordinates_are_not_deduplicated dedup does not over-collapse
test_different_coordinate_systems_are_not_deduplicated a normalised box and a pixel box are different claims
test_same_claim_from_two_specialists_is_kept_and_marked_corroborated agreement is recorded, not collapsed — the second specialist's name is kept as corroboration
test_unrelated_items_carry_no_corroboration_marker the marker is meaningful, not always-on
test_multiple_results_aggregate_without_loss the headline no-loss assertion
test_aggregation_does_not_double_count_a_specialists_own_evidence lossless ≠ duplicating

The corroboration tests are the direct expression of the docstring's "agreement between specialists must be recorded rather than collapsed with one name thrown away." This is a design decision encoded as a test: when two specialists agree, the collection records both names and marks the claim corroborated, rather than keeping one and discarding the other.

The cap tests pin that truncation is counted, not silent:

Test What it pins
test_zero_max_items_is_rejected an invalid cap is refused
test_cap_is_respected_and_the_drop_is_counted items past the cap are dropped and counted
test_no_truncation_flag_when_under_the_cap the truncation flag is truthful
test_default_cap_matches_the_config_value DEFAULT_MAX_ITEMS tracks the config
test_sources_are_reported_before_the_cap_is_applied which specialists contributed is recorded before truncation

That last test is a specific honesty property: if the cap drops a specialist's only item, the collection still records that the specialist contributed. Losslessness at the source level survives truncation at the item level.

The empty/degenerate cases are pinned too: test_empty_input_yields_an_empty_collection, test_a_result_with_no_evidence_contributes_nothing, test_accepts_a_single_result_without_wrapping, and the two argument-error tests test_both_results_and_evidence_is_rejected / test_neither_results_nor_evidence_is_rejected.

4.5 The confidence honesty rule

This is where the project's style guide (DOCS_STYLE_GUIDE.md §3) meets a test. The rule the engine must obey: an uncalibrated engine passes the raw score through and says it is uncalibrated; it never manufactures a fitted mapping. The tests that pin this:

Test What it pins
test_no_calibration_passes_raw_through_and_says_so the default is identity + an explicit "uncalibrated" label
test_uncalibrated_is_never_labelled_as_fitted the label cannot be silently upgraded
test_identity_temperature_is_reported_as_uncalibrated T = 1.0 is not calibration
test_engine_without_calibration_does_not_claim_calibration the engine-level claim
test_raw_outside_unit_range_is_clamped_not_rejected a robustness decision, pinned
test_non_finite_raw_is_refused NaN/Inf are refused, not clamped

This matters because the project's calibration genuinely made the metric worse (ECE 0.013755 → 0.014929, DOCS_STYLE_GUIDE.md §3) and is retained only because it is in the frozen config. The engine must therefore never present calibration as an improvement, and these tests are what enforce that at the code level. The METHOD_TEMPERATURE / METHOD_UNCALIBRATED constants are the two labels, and test_uncalibrated_is_never_labelled_as_fitted is the guard that the first is never used for the second.

The fitted path is tested separately and is deterministic: test_temperature_scaling_is_applied_when_fitted, test_temperature_above_one_softens_and_below_one_sharpens, test_temperature_scaling_is_monotonic, test_temperature_scaling_is_deterministic, test_endpoint_scores_do_not_produce_nan, test_invalid_temperatures_are_rejected, and test_specialist_degradation_survives_calibration / test_specialist_components_are_preserved_through_calibration.

4.6 Calibration artifact I/O

The engine reads a calibration artifact from disk, and the tests pin the failure modes explicitly rather than letting them degrade silently:

Test What it pins
test_artifact_round_trips_through_json write→read is lossless
test_nested_artifact_form_is_accepted both artifact shapes are accepted
test_missing_artifact_degrades_instead_of_failing a missing artifact is a degradation by default
test_missing_artifact_raises_when_required …but raises when the caller says it is required
test_malformed_artifact_is_not_silently_swallowed a corrupt artifact is an error, not a shrug
test_artifact_without_a_temperature_is_rejected the required field is required
test_load_calibration_respects_the_master_switch the config switch is honoured
test_load_calibration_reads_the_configured_file the configured path is the one read
test_load_calibration_without_a_file_configured_is_none no config ⇒ None, not a guess
test_engine_from_config_uses_the_configured_cap / ..._falls_back_to_the_default_cap the cap wiring

The pairing of degrades_instead_of_failing with raises_when_required is the pattern this repository uses throughout: a soft default for a convenience path, and a hard error for the path where silence would be a lie.

4.7 The scratch-root workaround, and why it is documented in the test file

The suite does not use pytest's built-in tmp_path. The module says why, and the reason is a sandbox property, not a preference:

"Scratch root, repo-local. The built-in tmp_path fixture cannot finalize under this machine's sandbox, so artifact tests use this instead. It is the same directory --basetemp=.pytest_tmp points pytest at, so nothing is written outside the workspace." — tests/unit/test_evidence_engine.py

SCRATCH_ROOT = Path(__file__).resolve().parents[2] / ".pytest_tmp"

And the module docstring carries the operational requirement:

"The --basetemp=.pytest_tmp flag is mandatory (see the project handoff): the default pytest temp root triggers a sandbox denial on this machine."

The scratch fixture is deliberately not cleaned up on teardown, and the docstring says why: "the sandbox's safe-delete guard rejects the recursive delete, and a test that writes into the repo's own ignored scratch directory is harmless. Each test gets a unique path so no state leaks between them." This is the same guard that causes §7's four test_safe_delete_shim failures — documented here as a design constraint rather than hidden as a quirk.

4.8 The public surface

test_module_public_surface_is_declared pins the exported names, and test_serialises_through_the_schema_serialiser pins that the collection serialises through the schema serialiser (not a bespoke one), so the wire form and the evidence form cannot drift. The shorthand and helpers are pinned by test_aggregate_evidence_shorthand_matches_the_engine, test_collection_helpers_filter_correctly, test_collection_summary_reports_the_observable_facts, test_digest_changes_when_a_claim_changes, test_digest_is_stable_across_regenerated_uuids, test_collection_is_iterable_and_sized.

test_digest_is_stable_across_regenerated_uuids is the one that makes §9 possible: because the digest ignores process-local uuids, a recorded run's evidence digest can be recomputed in a different process and still match. Without it, "recompute the verdict from the record" would be a claim you could not test.

4.9 What the evidence-engine suite does NOT establish

  • It does not establish that the specialists produce correct evidence — only that the aggregator is deterministic and lossless given whatever they produce.
  • It does not establish end-to-end reproducibility of a live run; that is §8's job.
  • It does not establish that the calibration is good — the tests pin that calibration is honestly labelled, not that it improves anything (it does not; ECE got worse).

5. The frontend live-wiring suite

5.1 What it proves, and why it reads source text

tests/unit/test_frontend_live_wiring.py is the regression net for the frontend↔orchestrator boundary. Its docstring states two things it proves, both of which were broken before the change:

"A. CORS. The deployment allowed exactly one origin (https://satquery.pages.dev), so the frontend could not be driven from a local development server at all: every request from http://localhost:8080 was refused with Disallowed CORS origin, and a developer had to deploy to Cloudflare to test a one-line JavaScript change. The fix adds explicit localhost origins while keeping production listed and keeping * rejected.

B. The task vocabulary. The page's router and the server's Task enum must agree. They are two independently written vocabularies in two languages, and nothing previously checked that a value the frontend would send is a value the server accepts. This module reads BOTH and compares them." — tests/unit/test_frontend_live_wiring.py

And it pins a third, subtler defect — one invisible to a validity check:

"It also pins the chang word-boundary defect: \bchang\b cannot match 'changed', so the page's own default question ('What changed here?') routed to vqa instead of the change path. That is invisible to any test that only checks 'is this a valid enum value' — both branches produce a valid value. The test therefore asserts the ROUTE, not just the validity."

That last sentence is the philosophical centre of the whole suite: assert the route, not the validity. A test that asks "is vqa a valid task?" passes for both the correct and the defective router, because both emit valid tasks. Only a test that asks "does this question route to this task?" can see the defect.

Why source-text parsing. The module reads the frontend source rather than importing it, and says so:

"the frontend is plain ES5 with no build step and no module exports, so there is no importable symbol. Parsing the source is the only way to assert the shipped bytes. The parsing is anchored on named tokens rather than line numbers, so it does not silently pass if the file is restructured."

The anchors are module-level path constants — REPO_ROOT, DEPLOY_RENDER, FRONTEND_JS, MISSION_JS, LIVE_JS, CORE_JS, MISSION_HTML — and the production origin is a single constant, PRODUCTION_ORIGIN = "https://satquery.pages.dev". The CORS half imports the orchestrator with a clean environment via _render_module(), because _allowed_origins reads os.environ at call time while _DEV_ORIGINS is built at import time — a distinction the module's helper comment records explicitly.

5.2 The measured result

106 passed (release/CURRENT_RELEASE_STATE.md:121; README.md §The test suites; docs/EVALUATION.md §7.1). The command is:

.venv/Scripts/python.exe -m pytest tests/unit/test_frontend_live_wiring.py -q     # 106 passed

This is a re-run-this-session figure, and it is the number this chapter quotes.

5.3 The fifteen test classes

The classes were read from the file. Each name is a claim:

# Class What it pins
1 TestTheCorsAllowlistKeepsProductionAndAddsDevelopment production origin always allowed, not duplicated, dev origins added, * rejected even when hidden in a list
2 TestTheCorsDecisionMatchesTheAllowlist the response-leg decision matches the allowlist (echo, no-headers, lookalike-host rejection)
3 TestTheFrontendSpeaksTheServersTaskVocabulary every mapped value is a real server task; the mapping covers every task the router can emit; specialist names are human labels, not tasks
4 TestTheDefaultQuestionReachesTheChangePath the page's own default question routes to the change path; the change stem matches its inflections; the temporal slot is required for a change question
5 TestALocationQuestionIsNotAChangeQuestion a "where" question reads as grounding, with one or two assets
6 TestTheArchitecturePolicyDoesNotReadALocationAsAChange the architecture sample routes to grounding; a place word in a "where" question is not a change
7 TestTheTaskRespectsThePairRequirement paired tasks substitute correctly when given one asset; single-asset tasks are never substituted
8 TestTheLiveClientUsesTheRealIngestionPath the client targets the orchestrator's routes; uploads send raw bytes with a derived content type; the infer body carries asset ids under assets; the client never sends a filesystem path
9 TestThePageLoadsTheLiveClient mission.html includes live JS before mission JS; the declared API base is an absolute origin, not the relative fallback, not the frontend origin; no fixture image is wired into the analysis path
10 TestTheLivePathKeepsTheHonestFailureContract a failure is attributed to the step that failed; the preview is kept for the no-file case
11 TestTheCaptionStopsCallingARealUploadIllustrative a successful run rewrites the leading word and the alt text; the failure path does not claim an analysis happened; a new file clears the previous run's claim
12 TestTheFrontendSendsOnlyTheAssetsTheTaskRequires a single-asset task with a pair uploaded uses only t1; the old wiring would have sent both; change/optical-SAR send the pair when present
13 TestDescriptiveQueriesRouteToCaption descriptive queries reach caption; a description is not mistaken for change; caption is single-asset
14 TestTheOpticalSarInputValidation a real optical+GeoTIFF SAR pair is ok; two plain photos warn; a missing SAR image is an error; the warning names the optical+radar expectation
15 TestTheServerErrorTranslation invalid request gets asset guidance; 422 gets malformed guidance; recoverable false gets restart guidance; the server's message and detail are preserved; a null error is handled; a recoverable transport code is not called unrecoverable

Two of these deserve a note because they encode counter-intuitive decisions:

TestTheFrontendSendsOnlyTheAssetsTheTaskRequires (class 12). The class contains test_the_old_wiring_would_have_sent_both_files_for_vqa and test_the_fix_sends_one_file_where_the_old_sent_two. This is a differential test: it asserts not just that the new code is right, but that the old code was wrong in the specific way claimed. That is the falsification discipline (§1.3) applied inside a suite — the fix is only meaningful if the defect it fixes is real, and the test proves the defect by asserting the old behaviour differs.

TestTheCaptionStopsCallingARealUploadIllustrative (class 11). This is a copy-honesty test. The page used to caption a real analysis "illustrative"; the suite asserts the leading word and the alt text are rewritten on success, and — crucially — that the failure path does not claim an analysis happened (test_the_failure_path_does_not_claim_an_analysis_happened). A UI that says "analysis complete" on a failure is a claim defect, and it is tested as one.

5.4 The historical count: 94 → 100 → 106

A reviewer may see a different number. docs/FINAL_DELIVERY_REPORT.md §7 records 94 passed for this file, while the current figure is 106. The difference is the suite growing after the delivery report was written: the regression tests added for the router fix took this file from 100 → 106 tests (DELIVERY_REPORT_2026-09-25.md, regression-tests note). This chapter quotes 106 and cites both figures so neither looks like a contradiction (docs/REPRODUCIBILITY.md §5.2).

5.5 What the frontend live-wiring suite does NOT establish

  • It does not run the browser. It asserts the shipped bytes of the frontend against the orchestrator module in-process. Whether the page actually loads and runs is §8's job (live validation).
  • It does not prove the router's quality — only that specific questions route to the specific tasks asserted. The router's residual misroutes ("What is the new runway?" → change) are real and recorded (docs/LIMITATIONS.md; LIVE_VALIDATION_POSTFIX.md), not hidden by a green suite.
  • It does not test the inference service. It tests the boundary the frontend speaks to, which is the orchestrator.

6. The doc/frontend suite

6.1 The five files, and the measured result

The "doc/frontend suite" is a named grouping of five files that together report 183 passed (docs/FINAL_DELIVERY_REPORT.md §7; docs/EVALUATION.md §7.2; docs/REPRODUCIBILITY.md §5.3). The five are:

File What it pins
tests/unit/test_frontend_guide_doc.py both documents agree on tasks, coordinate systems, evidence types, geometry shapes; no login screen; no secrets in the browser; the confidence-property trap; serialized analyses
tests/unit/test_frontend_live_wiring.py the frontend↔orchestrator boundary — see §5
tests/unit/test_api_contract_doc.py every JSON block parses and validates; no contract example fabricates an artifact URI; documented task/coordinate-system values are the real ones; every error code documented; the frozen config hash is recorded; the calibration caveat is recorded
tests/unit/test_runbook_doc.py the runbook names only public serving entrypoints; quoted artifact sizes match the real files; the capability list matches the registry; env vars are the architecture ones; undecided items declared undecided; no claim that a deployment happened; verified distinguished from designed
tests/unit/test_deploy_config.py the deploy manifest declares the non-registry marker; the validator reports no errors; the config-hash regression guard; the C-8 constraints (torch_compile false, cpu mode required, lazy load, single model cache)

These are documentation conformance tests: they fail if a doc drifts from the code it describes. They are the mechanism that keeps this release's documentation honest — the reason a doc claim can be trusted is that a test would go red if the claim stopped matching the code.

6.2 Why four of the five are doc-guards, and one is not

Four of the five (test_frontend_guide_doc, test_api_contract_doc, test_runbook_doc, test_deploy_config) are doc-guards, analysed individually in §3. The fifth (test_frontend_live_wiring) is not a doc-guard — it is a code-behaviour suite that happens to be grouped here because it reads frontend source text. The grouping is by audience (frontend + documentation) rather than by mechanism, and that is why the count 183 covers both.

6.3 The doc-guard guarantee, restated

The doc-guards guarantee example-level and value-level conformance, not prose truth (§3.7). The two module docstrings state the limit themselves:

"These tests validate the EXAMPLES, not the prose. They cannot detect a wrong sentence (e.g. a mis-stated latency budget)." — tests/unit/test_api_contract_doc.py

"Prose can still be wrong in ways these tests cannot see -- a mis-stated latency, a wrong reason for a decision." — tests/unit/test_frontend_guide_doc.py

test_runbook_doc.py reaches a different class of claim — numbers, names, and status honesty — and its most important assertion is that the document itself labels its unimplemented sections as unimplemented (§3.4). That is the guard against the exact overclaim the style guide forbids.

6.4 The config-hash regression guard

test_deploy_config.py carries the anti-drift move that makes it more than a validator: it does not re-declare the frozen config hash. It imports it:

"We import it rather than re-declare it, so the two cannot drift." — tests/unit/test_deploy_config.py

The authority is tests.test_config, and the value is the frozen hash 78f1e3700da15aa1 (DOCS_STYLE_GUIDE.md §3). A second copy of the hash anywhere in the tree would be a latent contradiction; the import makes drift structurally impossible.

The same file proves the manifest is INERT — "it is never merged into the config registry and cannot move Config.hash" — and that negative cases use in-memory dicts via validate_documents "so the real config files are never mutated." It also drives the validator CLI in a child interpreter to prove the module's own sys.path bootstrap makes core importable "WITHOUT pytest's pythonpath = . shortcut" (test_validator_cli_exits_zero_as_a_standalone_script).

6.5 What the doc/frontend suite does NOT establish

  • It does not establish that any deployment exists. test_runbook_doc.py explicitly asserts the opposite direction: that the document does not claim a deployment happened (§3.4).
  • It does not validate prose (§6.3).
  • It does not check that the code is correct — only that the docs match the code. If the code and the docs are wrong together, a doc-guard is green.

7. The full tests/unit result, and the 5–6 environmental failures

This is the section where the release's honesty discipline is most visible. Running the entire unit tree produces failures. They are reported, classified, and not reclassified as passes.

7.1 The result

Running python -m pytest tests/unit in the authoring sandbox produces 5–6 failures. Every one is environmental or ordering-related, not a regression in shipped code. The classification, exactly as docs/FINAL_DELIVERY_REPORT.md §7 records it:

# Failure Attribution Regression?
1–4 test_safe_delete_shim — 4 failures the sandbox's bulk-delete guard (Windows verbatim-path behaviour) No
5 one ordering flake in the router route test test ordering (passes in isolation) No
6 one stale adapter test (optical_sar absent when CROMA unshipped) stale test — CROMA is now shipped No

7.2 The evidence that they are not regressions

The evidence is the re-run. Re-running the affected files together gives 137 passed (docs/FINAL_DELIVERY_REPORT.md §7; docs/EVALUATION.md §7.3–7.4). The logic is stated plainly:

"if the failures were real regressions in the code under test, re-running those files together would still fail. They do not — which is what distinguishes an environmental/ordering failure from a regression."

This is the falsification discipline (§1.3) applied to a test result rather than to a guard: the claim "these are not regressions" is itself tested, by a re-run designed to fail if it were false.

7.3 The four test_safe_delete_shim failures, in detail

tests/unit/test_safe_delete_shim.py (338 lines, pytestmark = pytest.mark.unit) guards two defects in the WorkBuddy Windows safe-delete shim (cli/vendor/shim/sitecustomize.py) that produced phantom test failures in this repository. The module's docstring records both with measurements.

Defect 1 — a verbatim temp path was not recognised as a temp path.

"_path_for_compare returned normcase(realpath(abspath(path))), which preserves the Windows verbatim prefix \\?\. os.path.relpath cannot relate \\?\c:\... to c:\..., so _is_under_root returned False and the OS-temp exemption in _should_bypass_safe_delete did not apply."

Measured, before the fix:

plain    temp subdir  -> bypass True
verbatim temp subdir  -> bypass False      <-- the bug
non-temp subdir       -> bypass False

That is why routine pytest garbage-* collection — which walks \\?\C:\Users\...\Temp\pytest-of-anish\garbage-* — reached the bulk guard, tripped confirmRequired at 69 entries against a threshold of 50, and latched a rejection that then blocked every delete in the conversation.

Defect 2 — a successful delete was reported as a failure.

"_platform_trash raised whenever SHFileOperationW returned non-zero."

Measured on the authoring host with FOF_ALLOWUNDO|FOF_NOCONFIRMATION|FOF_NOERRORUI|FOF_SILENT, calling shell32 directly (the shim not involved):

series                                   non-temp rc      temp rc
4 regions x 3 interleaved rounds         2, 2, 2, ...     0, 0, 0
6 processes x 20 deletes                  2 (120/120)      0 (120/120)

The target is removed in every case and the user's Recycle Bin is populated and active (1,300+ $I/$R entries), so the return code is not a reliable "could not delete" signal. The module is explicit that the code is intermittent and that the precise Windows-internal trigger is not known:

"across pytest invocations the same non-temp delete sometimes returned 0, and in one run a temp delete returned non-zero. The precise Windows-internal trigger is NOT established and is not claimed here."

This is a place where this chapter writes UNKNOWN — not established from the available evidence, following the module's own honesty.

What the guard asserts (the WHAT IS ASSERTED block of the docstring):

  • A verbatim-prefixed path inside the OS temp root compares equal to its plain form and is exempted.
  • A verbatim-prefixed path outside the temp root is still guarded — "the fix normalises the prefix, it does not widen the exemption."
  • Deleting through a verbatim temp path succeeds end to end.
  • A non-zero shell code on a delete whose target is genuinely gone does not raise (stubbed shell, stubbed lexists — deterministic, deletes nothing).
  • A non-zero shell code on a delete whose target survives still raises — "Fail-closed is preserved; only the false positive is gone."
  • Against the real shell, a non-temp delete does not raise (behavioural confirmation, not the guard).

Why the guard is stub-based. The module states the reason, and it is a direct consequence of defect 2's intermittency:

"the guard for this defect stubs the shell: a test that depends on the real return code would sometimes pass against the broken shim. The stub tests fail on the original shim on every run."

That is the difference between a test that sometimes catches a defect and a test that always does.

The rule for extending the file is stated in the docstring and is worth reproducing, because it is the exact failure this file exists to prevent:

"Do not perform a real delete outside the OS temp root. … An earlier draft of the defect-2 test did exactly that, tripped confirmRequired at the threshold and latched a rejection that blocked every subsequent delete in the session — the very failure this file exists to prevent."

The module also skips cleanly when the shim is absent or disabled, via pytest.importorskip("sitecustomize", ...) plus skipif markers (needs_shim_helpers, windows_only). So on a normal user environment these tests skip; in the authoring sandbox they reach the guard and 4 of the 8 fail. The constants involved are VERBATIM = "\\\\?\\" and a NON_TEMP probe path that is "never created" — the predicates tested are pure.

7.4 The ordering flake

One router route test fails only when the full suite runs in a particular order, and passes in isolation. That is the definition of an ordering flake: the failure is a function of collection order, not of the code. docs/REPRODUCIBILITY.md §5.4 records it in the same row-set as the shim failures.

7.5 The stale adapter test

One test asserts optical_sar is absent when CROMA is unshipped. CROMA is now shipped (CURRENT_RELEASE_STATE.md §1 capability contract: optical_sar available: true), so the assertion's premise has changed. This is a stale test, not a code defect: the test encodes an assumption about the environment that the environment has outgrown.

7.6 A different environment state: the 2,237-passed run

docs/PHASE12_115_METRIC_COMPUTED.md §8 records, for 2026-09-22, a full unit suite result of 2,237 passed, 0 failed (622.68 s). That is a real, measured result in a different environment state (the safe-delete shim's guard state and the CROMA shipping status differ). It is recorded rather than suppressed, "because the two results are both true and the difference is exactly the environmental/ordering story" (docs/EVALUATION.md §7.3).

The same document records a correction that is itself an honesty lesson: an earlier revision claimed 2,179 passed computed as 2,171 + 8, "presented as a check but the total was never measured — it was inferred from a stale baseline." The measured figure was 2,237. This is the style guide's rule (DOCS_STYLE_GUIDE.md §1.3) in the wild: an inferred number presented as a measurement was corrected to the measured one.

7.7 The UNKNOWN: the collected count

The exact collected test count for the current-session full-suite run is UNKNOWN — not established from the available evidence. The record gives the failure classification and the 137-passed re-run, but not a collected total. Per-suite counts are known (106, 183, 137, 73, 51); the single collected total is not.

7.8 Why this is reported rather than hidden

The style guide's first rule is DO NOT FABRICATE, and its status rule is "a mixed result is never 'all work perfectly'" (DOCS_STYLE_GUIDE.md §1.7). A full-suite run with 5–6 failures is a mixed result. Presenting only the 106 and 183 would be exactly the "all work perfectly" overclaim the guide forbids. The correct presentation is the three-row table (§7.1) plus the re-run evidence (§7.2) plus the explicit UNKNOWN (§7.7) — which is what this section is.


8. The live-validation harness and its integrity discipline

8.1 What live validation is, and why it is separate from the suites

Every suite in §2–§7 runs in-process. docs/API_CONTRACT.md §8 states the boundary plainly: "no test in this repository dials a network address." A green tests/integration run is therefore not evidence about a deployment. Live validation is the separate activity that drives the deployed stack — the real Cloudflare Pages frontend, the real Render orchestrator, the real tunnel, the real Codespace — with a browser, and records what happened.

The record lives at .workbuddy-ai/scratch/live_validation/ — LIVE_VALIDATION_POSTFIX.md, the raw harness outputs (run_output.txt, run_final2.txt, run_final3.txt), the recomputed verdicts (results_final.json, results_pass3.json), and eight per-case screenshots (A1_vqa.png … B2_newairport.png).

8.2 The three passes, summarised

Pass HEAD Harness Result
1 ff46eba42b18 + d413d3672311 v1 (fill_input) 8/8
2 2d7ae53b482d v2 asserting 8/8
3 2d7ae53b482d v2 asserting (pre-discriminator-fix) 8/8 (recomputed)

24 live runs, 24 correct dispatches, 0 mock nodes, no run id repeated across passes. The trace fill was 94.4444 % on every case, and every case carries a real run_* id, the Hugging Face link in the DOM, and the capabilities → assets → infer call sequence all addressed to <backend-host> (two assets calls for the pair tasks).

The eight cases are Phase A (six regression cases: vqa, caption, grounding, change, change_vqa, optical_sar) and Phase B (the two defect cases: "Where are the built-up areas in this image?" and "Where is the new airport?", both expected to reach grounding with a single asset — the exact condition under which the old router collapsed to vqa and answered "River").

8.3 The CDP synthetic-key-event trap

This is the integrity lesson that shaped the harness, and it is the most transferable finding in this chapter.

The shape of the false pass. Pass 1's harness drove the query box with fill_input(), which types with real CDP key events. A re-run attempt failed on case 1 with run_id=0002, mock_nodes=9, answer="No answer yet", and only a capabilities call — the mock path. The root cause, measured directly:

"fill_input types with real CDP key events, and Chrome drops synthesized key events when the browser window does not hold OS focus; it has no assertion, so it clicked Run with the page's default query still in the box. Measured directly: with Chrome backgrounded, press_key("Z") left #qtext.value unchanged, while type_text("Q") (CDP Input.insertText, not focus-gated) inserted fine." — LIVE_VALIDATION_POSTFIX.md

The measurement is recorded in diag_focus.harness, which prints the value of #qtext after press_key("Z") and after type_text("Q") — showing the first is a no-op and the second works.

The generalisation. A harness that types into a form and then reads a result cannot tell the difference between:

  • "the query was entered, and the run used it", and
  • "the query was never entered, and the run used the page's default".

Both produce a result. The result of the second is a plausible answer to a different question, and a harness without an assertion records it as the verdict for the first. That is a silent false pass, and it is the reason the fix is not "use a different typing helper" but "assert the form state before dispatch".

8.4 The fix: deterministic query entry plus pre-dispatch assertions

The harness was rebuilt as run_all_postfix2.harness, whose header states the change:

"v2 sets the query deterministically (js value + input event, falling back to Input.insertText) and ASSERTS the form state before clicking, recording q_ok / obs_ok / t0_ok and a computed verdict per case."

The three assertions, and the failure each prevents:

Assertion Checks Failure it prevents
q_ok #qtext.value holds the intended query the silent-drop false pass
obs_ok #obsTail == 'ready' and one file on #fileInput an upload that did not land
t0_ok both frames present, for pair tasks a pair task run on one asset
no_mock_nodes mock_nodes == 0 the preview path being recorded as live

The rule, stated in the session handoff (HANDOFF_NEXT_AGENT.md §5.2) and reproduced in docs/architecture/09-frontend.md §10.3:

"Do NOT use fill_input() or press_key() to enter the query. They type with real CDP key events, which Chrome silently drops when the browser window does not hold OS focus … Use js() to set #qtext.value (plus input/change events) and/or type_text() (CDP Input.insertText, not focus-gated). Always assert the form state before clicking Run … or a no-op will be recorded as a pass."

The set_query() function in the harness implements exactly this: it focuses the box, clears it, tries type_text(q) (the non-focus-gated CDP path), reads the value back, and falls back to a direct js set with input/change events if the read-back does not match. It returns the value actually read back, and q_ok = (qgot == c["q"]). The assertion is on the read-back, not on the attempt.

obs_ok's second half is a real DOM fact: #obsTail reads none in the markup (mission.html:67) and the live driver flips it to ready on upload, so obs_ok proves the upload landed rather than merely that a click happened.

8.5 The three discriminators that proved the earlier 8/8 was clean

The response to a silent-false-pass risk is not to re-run and hope; it is to check whether the earlier result was contaminated, using evidence the harness recorded. The report does exactly that:

"I then checked whether the earlier 8/8 run was infected by the same silent failure. It was not:

  • its recorded intents are query-specific — A1 reads taskvqa…temporalnone, whereas the default query "What changed here?" would read taskchange…temporalrequired (exactly what the failed run showed);
  • its answers embed the query text — e.g. [grounding] Located 6 candidate region(s) for 'Where are the built-up areas in this image?';
  • A6 required two files (optical 4/12 + SAR 2/2 channels), which only the uploaded pair supplies.

So the 8/8 result is a valid measurement." — DELIVERY_REPORT_2026-09-25.md §3

The three discriminators generalise into a reusable checklist:

Discriminator What it proves
the recorded intent is query-specific the query reached the router
the answer embeds the query text the server received the intended query
a case requires an artefact only the setup supplies the setup really happened

This is why the harness records the full per-case payload — intent, answer, files_t1, files_t0, obs_tail, mock_nodes, api_calls — rather than just a pass/fail bit. The bit can be wrong; the raw record can be re-examined.

8.6 The two discriminator bugs (and why they were false failures, not false passes)

Two harness bugs were found and fixed, both causing false failures — the safe direction:

  1. The answer [task] tag exists only for region tasks. vqa/caption answers are bare ("Grassland", prose), so a tag-based verdict heuristic reported them as failures. The fix computes the dispatched task as answer_tag when present, else the intent's reading.
  2. The intent panel is a concatenated string. task([a-z_]+) must be matched non-greedily up to modality; a greedy match swallows the whole string. The fix is re.search(r"task([a-z_]+?)modality", intent).

The harness header states the asymmetry: "Two further harness bugs were found and fixed, both causing false failures." A false failure costs a re-run; a false pass costs the integrity of the result. The harness was built so that its bugs err toward false failure, and §9 makes even the false-failure case recoverable without a re-run.

8.7 Liveness traps discovered during the runs

Two operational traps are recorded because they look exactly like stalls:

  • browser-use block-buffers stdout even when redirected to a file, so the output file stays at 0 bytes until the process exits — indistinguishable from a stall. The harness now forces line-buffering (sys.stdout.reconfigure(line_buffering=True)) and prints CASE_START <id> per case, so the file grows case by case.
  • grep block-buffers when piped, so piping the harness through grep swallows all output if the pipeline is killed. Redirect to a file instead.

The harness also records ALL_DONE as a final marker, so a truncated run is detectable from the output file alone.

8.8 What the live validation does NOT establish

  • It is not the independent audit. LIVE_VALIDATION_POSTFIX.md opens by saying so: "This is a re-validation, not the independent audit: the prior independent audit is ../2026-09-25-00-37-41/live_validation/INDEPENDENT_LIVE_VALIDATION.md."
  • It does not establish model quality. The two verdicts are kept separate: Deployment/integration: PASS and Model quality: MIXED (caption and grounding meaningful; change/change_vqa plausible; VQA weak-but-related — A1 answers "Grassland"; optical-SAR returns a bare class index class_18, not a human label).
  • It does not establish load behaviour, latency ceilings, or cold-start times. It is eight cases, run three times.
  • It does not clear B-07: transient tunnel-agent gaps remain OPEN, and the patch is prepared, NOT deployed (CURRENT_RELEASE_STATE.md §6).

9. Verdict-independent evidence: recomputing from raw records

9.1 The property

The live-validation harness has a verdict field. The project's design decision is that the verdict must not be the primary record. The recorded evidence — q_ok, obs_ok, t0_ok, run_id, mock_nodes, intent, answer, api_calls — is recorded independently of the verdict computation, so a harness bug in the verdict logic can never silently turn a real failure into a pass, and can be corrected without a re-run.

9.2 The recomputation script

.workbuddy-ai/scratch/recompute_verdicts.py implements the property. Its docstring states the case that motivated it:

"The v2 harness recorded, per case: q_ok, obs_ok, t0_ok, run_id, mock_nodes, intent, answer, api_calls -- all independent of the verdict computation. Its verdict field was wrong for vqa/caption because the tag heuristic assumed every answer starts with "[task]"; vqa/caption answers are bare. This recomputes the verdict from the intent's task<name> slot (primary) plus the answer tag (cross-check only when present)."

The script is read-only over the recorded run and writes results_final.json. The core of it:

mi = re.search(r"task([a-z_]+?)modality", intent)
intent_task = mi.group(1) if mi else ""
tag = answer[1:answer.find("]")] if (answer.startswith("[") and "]" in answer) else ""
expect = d["expect"]

# The DISPATCHED task is what the server actually ran. The answer carries a "[task]"
# prefix for the region tasks (grounding/change/change_vqa/optical_sar); vqa and
# caption answers are bare, so fall back to the intent's reading there.
# NOTE: A5 is a legitimate case where the two differ -- the router READS "change" but
# dispatches "change_vqa" (the documented quantifier upgrade). That is not a failure.
dispatched = tag if tag else intent_task

checks = {
    "query_in_box": bool(d.get("q_ok")),
    "asset_attached": bool(d.get("obs_ok")),
    "pair_attached": bool(d.get("t0_ok")),
    "live_run": str(d.get("run_id", "")).startswith("run_"),
    "no_mock_nodes": d.get("mock_nodes") == 0,
    "task_dispatched": dispatched == expect,
}
verdict = "PASS" if all(checks.values()) else "FAIL"

Every input to the verdict is a recorded observation; the verdict is a pure function of them. The script also records read_vs_dispatched_differs — so the one legitimate read≠dispatch case (A5, the documented quantifier upgrade) is flagged, not failed.

9.3 Pass 3 is the proof of the property

Pass 3 was launched with the harness build that still carried the two discriminator bugs (§8.6), so its raw output says SUMMARY 0/8. The verdicts were recomputed from the recorded evidence by recompute_verdicts.py and are 8/8 PASS (results_pass3.json). The record states the significance:

"That is the intended safety property — the recorded evidence (run id, mock_nodes, intent, answer, assertions) is independent of the verdict computation, so a harness bug can never silently turn a real failure into a pass, and never costs a re-run to correct." — LIVE_VALIDATION_POSTFIX.md

This is the strongest form of the claim: the property was not merely designed, it was exercised. A run with a broken verdict computation was corrected post hoc from its own raw record, without touching the deployment or re-running the browser.

9.4 The generalisation

The pattern is: record observations, compute verdicts separately, keep the observations. It is the same shape as the evidence engine's digest (§4.8): the record is content-addressed and process-independent, and the interpretation is a function over it. Both make the same guarantee — that a conclusion can be re-derived from the evidence rather than trusted as an assertion.


10. How to run the tests

10.1 The interpreter

Use the repository virtual environment, not a bare python. On Windows:

.venv/Scripts/python.exe -m pytest <target> -q

docs/REPRODUCIBILITY.md §10.2 records this as the invocation that reproduces the published results. On Windows, pytest must be given -p no:cacheprovider because the sandbox refuses .pytest_cache writes (docs/BACKEND_DEPLOYMENT_RUNBOOK.md §0).

10.2 Run targeted files, not the whole tree

A full tests/unit run trips the sandbox's bulk-delete guard (the 4× test_safe_delete_shim failures of §7.3). The workaround is to run targeted suites:

.venv/Scripts/python.exe -m pytest tests/unit/test_frontend_live_wiring.py -q     # 106 passed
.venv/Scripts/python.exe -m pytest tests/unit/test_evidence_engine.py --basetemp=.pytest_tmp -q
.venv/Scripts/python.exe -m pytest tests/unit/test_deploy_config.py -p no:cacheprovider -q
.venv/Scripts/python.exe -m pytest tests/unit/test_api_contract_doc.py tests/unit/test_frontend_guide_doc.py -p no:cacheprovider -q

The runbook notes that multi-suite invocations in one command have been refused by the environment before; single suites are reliable (docs/BACKEND_DEPLOYMENT_RUNBOOK.md §2.6). The --basetemp flag is mandatory for the evidence-engine suite (§4.7): the default pytest temp root triggers a sandbox denial on this machine, and the module's scratch fixture writes to a repo-local .pytest_tmp instead.

10.3 pytest.ini

The configuration is four lines plus the marker registrations:

[pytest]
testpaths = tests
pythonpath = .
addopts = -q --tb=short

(pytest.ini:1-4). pythonpath = . is what makes from core.config import ... resolve from the repo root without an installed package — and test_deploy_config.py deliberately proves the module also bootstraps sys.path itself, "WITHOUT pytest's pythonpath = . shortcut" (§6.4). The markers are registered at pytest.ini:23-27 (§0.1), and filterwarnings silences three warning classes including rasterio.errors.NotGeoreferencedWarning (pytest.ini:5-8).

10.4 The output-buffering trap

Do not pipe pytest through grep in the authoring sandbox — output is block-buffered and a killed pipeline swallows it. Redirect to a file instead (docs/REPRODUCIBILITY.md §5.6). The same trap applies to the live-validation harness (§8.7).

10.5 The expected results

Suite Command Expected Status
Frontend live-wiring pytest tests/unit/test_frontend_live_wiring.py -q 106 passed VERIFIED
Doc/frontend suite pytest on the 5 doc/frontend files 183 passed VERIFIED
Evidence engine pytest tests/unit/test_evidence_engine.py --basetemp=.pytest_tmp 73 passed VERIFIED
Gateway policy pytest tests/unit/test_gateway_policy.py 51 passed VERIFIED
Full unit suite pytest tests/unit 5–6 environmental/ordering failures, rest pass; a re-run passes 137 MEASURED

(release/repo/README.md §The test suites; docs/PHASE19_FINAL_HARDENING.md §6; docs/REPRODUCIBILITY.md §5.1.) The precise full-suite collected count is UNKNOWN — not established from the available evidence (§7.7).


11. What is NOT tested

A green suite is not evidence about a property nobody wrote an assertion for (DOCS_STYLE_GUIDE.md §1.6; §0.3). This section lists the properties the release does not have test evidence for, so that no reader mistakes coverage for completeness.

Property State Why / what exists instead
Continuous integration does not exist there is no .github/ directory and no CI workflow in the repository. Every test result in this chapter was produced by a manual run in the authoring sandbox. Nothing runs the suites automatically on push.
End-to-end system benchmark NOT RUN (none exists) CURRENT_RELEASE_STATE.md §4: *"System-level end-to-end benchmark — NOT RUN (none exists)."* No system-level accuracy is claimed anywhere in the release.
Load / soak / concurrency NOT RUN no test exercises sustained load, concurrent users, or a long-running process. The per-IP rate limiter is a fairness control, not a security or capacity control (docs/DEPLOYMENT_ARCHITECTURE.md §5.2).
Adversarial / fuzz testing NOT RUN no fuzzing or adversarial-input suite. Input validation is tested by example (§2.2.9), not by search.
Cross-dataset generalization NOT RUN the measured metrics are per-dataset (LEVIR-CD-256, VRSBench, held-out test sets); no test measures transfer to an unseen dataset.
Browser / device matrix NOT RUN live validation ran in headed Chromium only. No Firefox/Safari/WebKit, no mobile, no accessibility audit.
Deployment verification in CI NOT RUN the runbook guard asserts the document does not claim a deployment happened (§3.4). Deployment reachability is established only by the manual live validation (§8).
Router test split NOT RUN CURRENT_RELEASE_STATE.md §4: router accuracy 0.965116 is validation, ungated, n = 86; "the test split was NOT RUN" (DOCS_STYLE_GUIDE.md §3).
Calibration as an improvement not established (and measured the other way) ECE went 0.013755 → 0.014929 — worse. The engine suite pins that calibration is honestly labelled (§4.5), not that it helps.
Optical-SAR / change-VQA acceptance OPEN the rulings are OPEN (CURRENT_RELEASE_STATE.md §4; DOCS_STYLE_GUIDE.md §3). A suite passing is not a model acceptance.
VLM adapter acceptance REJECTED metrics USABLE_VERIFIED (exact_match 0.963) but status ACCEPTANCE-REJECTED. USABLE ≠ ACCEPTED (DOCS_STYLE_GUIDE.md §3).
Tunnel reliability (B-07) OPEN transient tunnel-agent gaps; patch prepared, NOT deployed. No test can cover a defect that is not fixed in production.
The inference service's internals UNKNOWN the tunnel-agent source is not in the monorepo; docs/architecture/02-deployment-topology.md §3.3/§9.4 documents the topology as a measured fact, but the agent's internals are UNKNOWN — not established from the available evidence.
Credential write permission (HF) UNKNOWN CURRENT_RELEASE_STATE.md §7: HF token identity thundercode, role fineGrained; "write permission not yet proven."

Two of these deserve emphasis:

  • No CI is the largest gap. Every count in this chapter is a manual measurement. A reader should not read "106 passed" as "the suite passes on every commit" — it means "it passed when it was run, in the recorded environment." There is no automation that would catch a regression on push.
  • No end-to-end benchmark is the second. The individual task metrics are artifact-backed and measured, but there is no measurement of the system — router + specialist + envelope + frontend — on a held-out end-to-end corpus. The live validation (§8) proves the pipeline runs and dispatches correctly; it does not measure system accuracy.

12. Open items and evidence index

12.1 Open / blocked / not-run items for this chapter's topic

Following the style guide's requirement that every doc end with an explicit list (DOCS_STYLE_GUIDE.md §4):

Item State Note
B-07 — transient tunnel-agent gaps OPEN patch prepared, NOT deployed; root shape measured (in auto mode a tunnel timeout falls through to the forward path, burning wake_timeout_s=120 on a 302 ≈ 249 s)
B-02 — codespace_name trailing \n OPEN (cosmetic) confirmed still live; the wake path strips it
No CI does not exist no .github/ directory; all results are manual
System-level end-to-end benchmark NOT RUN (none exists) no system-level accuracy claimed
Router test split NOT RUN val-only, n = 86
Full-suite collected test count UNKNOWN per-suite counts known; the single collected total is not
test_safe_delete_shim failures environmental sandbox delete-guard; not regressions; precise Windows trigger UNKNOWN
Ordering flake environmental passes in isolation
Stale adapter test stale asserts CROMA unshipped; CROMA is shipped
Live validation MEASURED 3 passes × 8 cases, 8/8 each, 24 runs, 0 mock nodes, trace fill 94.4444 %
Optical-SAR / change-VQA rulings OPEN 0.931 / macro-F1 0.434161; 0.697626 / 0.378373
VLM adapter acceptance REJECTED metrics usable; acceptance rejected
Calibration not an improvement ECE worse (0.013755 → 0.014929); retained only because it is in the frozen config

12.2 Where the evidence lives

What Where
Test configuration and markers pytest.ini
Frontend live-wiring suite tests/unit/test_frontend_live_wiring.py — 106 passed
Evidence-engine suite tests/unit/test_evidence_engine.py — 73 passed
Safe-delete shim guard tests/unit/test_safe_delete_shim.py — 338 lines; skips cleanly off-sandbox
Doc-guards tests/unit/test_api_contract_doc.py, test_frontend_guide_doc.py, test_runbook_doc.py, test_deploy_config.py, test_step7_api_contract_conformance.py
Integration suite scope tests/integration/test_step7_backend_chain.py, test_step7_r02_serving.py; docs/ITEM5_INTEGRATION_SUITE_SCOPE.md
The suite results, reported in full docs/EVALUATION.md §7.1–7.5; docs/REPRODUCIBILITY.md §5; release/repo/README.md §The test suites
The full-suite failure classification docs/FINAL_DELIVERY_REPORT.md §7; docs/PHASE12_115_METRIC_COMPUTED.md §8 (the 2,237-passed run)
Live-validation record .workbuddy-ai/scratch/live_validation/LIVE_VALIDATION_POSTFIX.md
Live-validation raw output .workbuddy-ai/scratch/live_validation/run_output.txt, run_final2.txt, run_final3.txt
Live-validation screenshots .workbuddy-ai/scratch/live_validation/A1_vqa.png … B2_newairport.png (8)
Live-validation harness (v1) .workbuddy-ai/scratch/run_all_postfix.harness
Live-validation harness (v2, asserting) .workbuddy-ai/scratch/run_all_postfix2.harness
CDP focus diagnostic .workbuddy-ai/scratch/diag_focus.harness
Verdict recomputation .workbuddy-ai/scratch/recompute_verdicts.py
Prior independent audit ../2026-09-25-00-37-41/live_validation/INDEPENDENT_LIVE_VALIDATION.md
Release-state inventory release/CURRENT_RELEASE_STATE.md

12.3 Cross-references