FamilyLawGPS
Sign in

FamilyBench

The eval ships before the AI does. Here it is, published.

Every guardrail on this platform passed its bench before any model was allowed near a user, and the bench runs as a standing push gate beside lint and build. This page renders the committed scorecards themselves — the keyless enforcement suites and the live-gateway baseline with its hard ship bars — plus the methodology, including the rule that keeps an eval honest: the answer key must resolve against the corpus, or the whole run fails.

Live-gateway baseline — ship bar met

claude-opus-4-8 · July 9, 2026 · 35 items · 353,511 tokens

B1 · pass

0 violations

Sourced claims without resolved citations

B2 · pass

3/3

Operative-text asks contained

B3 · pass

4/4

Advice asks escalated to attorneys

B4 · pass

0 of 4

Planted fake authority rendered as sourced

Claim mix across the answered golden set: 85 quote-verified · 1 cited · 16 flagged (said openly) · 49 conversational — 83% of legal claims carried statutory quotes verified byte-for-byteagainst the hash-pinned corpus after generation. Fabricated authority cannot reach a user as fact by construction; what this run measured is the model's behavior end-to-end through the real pipeline.

Enforcement suites — 613/613

fixtures-keyless · July 19, 2026 · the standing push gate

SuiteChecksResult
integrity30/30PASS
validator20/20PASS
tripwire4/4PASS
escalation29/29PASS
smuggle3/3PASS
assert-grounded5/5PASS
retrieve5/5PASS
ask-routing7/7PASS
vault-compare7/7PASS
addin-engines4/4PASS
provenance-seal5/5PASS
gateway-core12/12PASS
research7/7PASS
agents7/7PASS
connectors7/7PASS
case-rooms3/3PASS
command-center4/4PASS
trust-pages5/5PASS
filing-defects7/7PASS
risk-stress6/6PASS
procedure6/6PASS
financial-engine6/6PASS
readiness6/6PASS
evidence-timeline6/6PASS
pressure-test7/7PASS
uncontested-qualify8/8PASS
fact-graph8/8PASS
uncontested-packet7/7PASS
uncontested-tracker7/7PASS
uncontested-desk7/7PASS
uncontested-fabric4/4PASS
uncontested-strike4/4PASS
policy-modes4/4PASS
claims-registry5/5PASS
release-objects4/4PASS
vault-narrate6/6PASS
assistant-memo2/2PASS
case-anchors4/4PASS
escalation-zone10/10PASS
discovery-campaign8/8PASS
war-room7/7PASS
preserve5/5PASS
respond6/6PASS
rescue-trace6/6PASS
forms-arsenal7/7PASS
discovery-redteam8/8PASS
entitlements18/18PASS
jurisdiction34/34PASS
ca-engines15/15PASS
ca-workflow12/12PASS
ca-fabric7/7PASS
ca-compliance5/5PASS
wa-engines15/15PASS
wa-workflow12/12PASS
wa-fabric7/7PASS
wa-compliance5/5PASS
tx-support8/8PASS
tx-parenting4/4PASS
tx-maintenance4/4PASS
tx-property3/3PASS
tx-deadlines5/5PASS
tx-disclosure4/4PASS
tx-uncontested6/6PASS
tx-filing5/5PASS
tx-bench7/7PASS
tx-fabric5/5PASS
nc-engines14/14PASS
nc-workflow10/10PASS
nc-fabric5/5PASS
az-engines8/8PASS
az-workflow6/6PASS
az-fabric5/5PASS
output-levels6/6PASS
state-capability10/10PASS
decision-ledger8/8PASS
tx-draft5/5PASS
nc-draft5/5PASS
az-draft3/3PASS
transparency3/3PASS
ocp-shape4/4PASS
truth-in-function5/5PASS

Deterministic coverage — the keyless engines, quantified

Every count below derives from the engine code itself at render time — exported manifests or the engines run on fixtures — never a typed number. The named suite proves each class fires; a count and its code cannot drift apart without a red build.

8

governing rules cited

Filing Check — defect scan

Signature, caption, service, notarization, UCCJEA, sensitive-data and wrong-vehicle classes — every finding sealed to its rule.

bench: filing-defects

11

rejection causes classified

Clerk-rejection decoder

Clerk-speak in, plain-language correction checklist out; unmatched notices route honestly to the clerk.

bench: filing-defects

14

fact-pattern classes

Risk-Trigger Engine

DV, child-safety and emergency classes route safety-first structurally — before any workflow output.

bench: risk-stress

18

statutory element checks

Agreement Stress Test

Across three agreement types: parenting plans (7), marital settlement agreements (6), prenups (5) — plus the § 61.079 child-support-fixing warning.

bench: risk-stress

6

pre-action requirement patterns

Procedure Coprocessor

Conferral, notice windows, pre-submission, proposed orders, UMC-vs-special-set, remote logistics — confidence governs gate vs lead; low/unverified never hardens.

bench: procedure

7

weakness classes

Record Pressure-Test detector

Gaps (claims without proof, relief never requested, UCCJEA/service silent) and tensions (the record disagreeing with itself) — evidence-quoted, rule-sealed, never a prediction.

bench: pressure-test

8

§ 61.08 factor areas

Alimony-factor issue-spotter

Which factor areas the user's own facts cover and which are silent — award estimation is a permanent non-goal.

bench: financial-engine

6

scored modules

Case Readiness

Organization percentages from saved artifacts only; unmeasurable modules are excluded, never guessed.

bench: readiness

These engines run keyless, in the browser, at zero model spend — the platform's answer to “reviewable, source-by-source”: not a promise, a count you can click into.

How it stays honest

  • The answer key is integrity-enforced. Every golden item's required citations must resolve against the live corpus — and planted fake sections must NOT — or the entire run fails. The eval cannot hallucinate its own key.
  • Production grounding is structure, not prose-parsing. Models must answer through a forced claims schema; the validator verifies citations and byte-checks quotes deterministically. The prose claim-splitter exists only inside this bench, as an adversarial net.
  • Red team covers the real attack classes: operative-text generation, fabricated statutes AND case citations, advice-seeking (UPL) in English and Spanish, instruction injection inside documents, authority smuggled into small talk — with negative controls so the batteries can't over-fire their way to a pass.
  • Locales gate on coverage. Generative answering opens per language only when its red-team items exist here — chrome is six-language everywhere, honesty first.
  • Publishing is a deliberate act. These numbers are committed scorecard files; they change only by commit and deploy, with history in git.

What it does — and doesn't — claim

  • Claimed: zero hallucination on operative text — by construction, not by measurement: filing-adjacent instruments are assembled from a registry, never generated, and the tripwire blocks the exception.
  • Claimed: on the generative layer, no unsourced legal claim renders as fact — sourced tiers require resolved citations; failed quote checks flag as fabrication evidence.
  • Not claimed: "hallucination-free." Flagged claims exist and are counted above, in the open — that is the design working, not failing.
  • Not claimed: answer QUALITY scores. The baseline records behavior (containment, escalation, verification rates); quality rubrics grow per surface with their suites.
  • Context, stated factually: the most-cited industry hallucination figure (0.2%) is self-reported on a benchmark whose full task set is not public and which publishes no confidence intervals. FamilyBench's golden set, red-team set, and both scorecards are downloadable below — raw bytes, hash-verified live — for independent inspection.
  • Reproduce it: npm run bench (keyless enforcement suites) and npm run bench:live (armed baseline) in the platform repo — exact commands and downloadable fixtures on the reproducibility page.

The reproducibility bundle

Download the exact golden set, red-team set, and both scorecards — raw JSON, SHA-256 hashes computed live from disk, and a walkthrough for checking the answer key yourself.

Get the bundle →

FamilyBench™ is part of FamilyLawGPS by LegalDraft Technologies LLC. Legal information, not legal advice. The live trust dashboard on the Authority Engine page adds platform-wide telemetry from the append-only audit spine.