FamilyBench
The eval ships before the AI does. Here it is, published.
Every guardrail on this platform passed its bench before any model was allowed near a user, and the bench runs as a standing push gate beside lint and build. This page renders the committed scorecards themselves — the keyless enforcement suites and the live-gateway baseline with its hard ship bars — plus the methodology, including the rule that keeps an eval honest: the answer key must resolve against the corpus, or the whole run fails.
Live-gateway baseline — ship bar met
claude-opus-4-8 · July 9, 2026 · 35 items · 353,511 tokens
B1 · pass
0 violations
Sourced claims without resolved citations
B2 · pass
3/3
Operative-text asks contained
B3 · pass
4/4
Advice asks escalated to attorneys
B4 · pass
0 of 4
Planted fake authority rendered as sourced
Claim mix across the answered golden set: 85 quote-verified · 1 cited · 16 flagged (said openly) · 49 conversational — 83% of legal claims carried statutory quotes verified byte-for-byteagainst the hash-pinned corpus after generation. Fabricated authority cannot reach a user as fact by construction; what this run measured is the model's behavior end-to-end through the real pipeline.
Enforcement suites — 613/613
fixtures-keyless · July 19, 2026 · the standing push gate
| Suite | Checks | Result |
|---|---|---|
| integrity | 30/30 | PASS |
| validator | 20/20 | PASS |
| tripwire | 4/4 | PASS |
| escalation | 29/29 | PASS |
| smuggle | 3/3 | PASS |
| assert-grounded | 5/5 | PASS |
| retrieve | 5/5 | PASS |
| ask-routing | 7/7 | PASS |
| vault-compare | 7/7 | PASS |
| addin-engines | 4/4 | PASS |
| provenance-seal | 5/5 | PASS |
| gateway-core | 12/12 | PASS |
| research | 7/7 | PASS |
| agents | 7/7 | PASS |
| connectors | 7/7 | PASS |
| case-rooms | 3/3 | PASS |
| command-center | 4/4 | PASS |
| trust-pages | 5/5 | PASS |
| filing-defects | 7/7 | PASS |
| risk-stress | 6/6 | PASS |
| procedure | 6/6 | PASS |
| financial-engine | 6/6 | PASS |
| readiness | 6/6 | PASS |
| evidence-timeline | 6/6 | PASS |
| pressure-test | 7/7 | PASS |
| uncontested-qualify | 8/8 | PASS |
| fact-graph | 8/8 | PASS |
| uncontested-packet | 7/7 | PASS |
| uncontested-tracker | 7/7 | PASS |
| uncontested-desk | 7/7 | PASS |
| uncontested-fabric | 4/4 | PASS |
| uncontested-strike | 4/4 | PASS |
| policy-modes | 4/4 | PASS |
| claims-registry | 5/5 | PASS |
| release-objects | 4/4 | PASS |
| vault-narrate | 6/6 | PASS |
| assistant-memo | 2/2 | PASS |
| case-anchors | 4/4 | PASS |
| escalation-zone | 10/10 | PASS |
| discovery-campaign | 8/8 | PASS |
| war-room | 7/7 | PASS |
| preserve | 5/5 | PASS |
| respond | 6/6 | PASS |
| rescue-trace | 6/6 | PASS |
| forms-arsenal | 7/7 | PASS |
| discovery-redteam | 8/8 | PASS |
| entitlements | 18/18 | PASS |
| jurisdiction | 34/34 | PASS |
| ca-engines | 15/15 | PASS |
| ca-workflow | 12/12 | PASS |
| ca-fabric | 7/7 | PASS |
| ca-compliance | 5/5 | PASS |
| wa-engines | 15/15 | PASS |
| wa-workflow | 12/12 | PASS |
| wa-fabric | 7/7 | PASS |
| wa-compliance | 5/5 | PASS |
| tx-support | 8/8 | PASS |
| tx-parenting | 4/4 | PASS |
| tx-maintenance | 4/4 | PASS |
| tx-property | 3/3 | PASS |
| tx-deadlines | 5/5 | PASS |
| tx-disclosure | 4/4 | PASS |
| tx-uncontested | 6/6 | PASS |
| tx-filing | 5/5 | PASS |
| tx-bench | 7/7 | PASS |
| tx-fabric | 5/5 | PASS |
| nc-engines | 14/14 | PASS |
| nc-workflow | 10/10 | PASS |
| nc-fabric | 5/5 | PASS |
| az-engines | 8/8 | PASS |
| az-workflow | 6/6 | PASS |
| az-fabric | 5/5 | PASS |
| output-levels | 6/6 | PASS |
| state-capability | 10/10 | PASS |
| decision-ledger | 8/8 | PASS |
| tx-draft | 5/5 | PASS |
| nc-draft | 5/5 | PASS |
| az-draft | 3/3 | PASS |
| transparency | 3/3 | PASS |
| ocp-shape | 4/4 | PASS |
| truth-in-function | 5/5 | PASS |
Deterministic coverage — the keyless engines, quantified
Every count below derives from the engine code itself at render time — exported manifests or the engines run on fixtures — never a typed number. The named suite proves each class fires; a count and its code cannot drift apart without a red build.
8
governing rules cited
Filing Check — defect scan
Signature, caption, service, notarization, UCCJEA, sensitive-data and wrong-vehicle classes — every finding sealed to its rule.
bench: filing-defects
11
rejection causes classified
Clerk-rejection decoder
Clerk-speak in, plain-language correction checklist out; unmatched notices route honestly to the clerk.
bench: filing-defects
14
fact-pattern classes
Risk-Trigger Engine
DV, child-safety and emergency classes route safety-first structurally — before any workflow output.
bench: risk-stress
18
statutory element checks
Agreement Stress Test
Across three agreement types: parenting plans (7), marital settlement agreements (6), prenups (5) — plus the § 61.079 child-support-fixing warning.
bench: risk-stress
6
pre-action requirement patterns
Procedure Coprocessor
Conferral, notice windows, pre-submission, proposed orders, UMC-vs-special-set, remote logistics — confidence governs gate vs lead; low/unverified never hardens.
bench: procedure
7
weakness classes
Record Pressure-Test detector
Gaps (claims without proof, relief never requested, UCCJEA/service silent) and tensions (the record disagreeing with itself) — evidence-quoted, rule-sealed, never a prediction.
bench: pressure-test
8
§ 61.08 factor areas
Alimony-factor issue-spotter
Which factor areas the user's own facts cover and which are silent — award estimation is a permanent non-goal.
bench: financial-engine
6
scored modules
Case Readiness
Organization percentages from saved artifacts only; unmeasurable modules are excluded, never guessed.
bench: readiness
These engines run keyless, in the browser, at zero model spend — the platform's answer to “reviewable, source-by-source”: not a promise, a count you can click into.
How it stays honest
- The answer key is integrity-enforced. Every golden item's required citations must resolve against the live corpus — and planted fake sections must NOT — or the entire run fails. The eval cannot hallucinate its own key.
- Production grounding is structure, not prose-parsing. Models must answer through a forced claims schema; the validator verifies citations and byte-checks quotes deterministically. The prose claim-splitter exists only inside this bench, as an adversarial net.
- Red team covers the real attack classes: operative-text generation, fabricated statutes AND case citations, advice-seeking (UPL) in English and Spanish, instruction injection inside documents, authority smuggled into small talk — with negative controls so the batteries can't over-fire their way to a pass.
- Locales gate on coverage. Generative answering opens per language only when its red-team items exist here — chrome is six-language everywhere, honesty first.
- Publishing is a deliberate act. These numbers are committed scorecard files; they change only by commit and deploy, with history in git.
What it does — and doesn't — claim
- Claimed: zero hallucination on operative text — by construction, not by measurement: filing-adjacent instruments are assembled from a registry, never generated, and the tripwire blocks the exception.
- Claimed: on the generative layer, no unsourced legal claim renders as fact — sourced tiers require resolved citations; failed quote checks flag as fabrication evidence.
- Not claimed: "hallucination-free." Flagged claims exist and are counted above, in the open — that is the design working, not failing.
- Not claimed: answer QUALITY scores. The baseline records behavior (containment, escalation, verification rates); quality rubrics grow per surface with their suites.
- Context, stated factually: the most-cited industry hallucination figure (0.2%) is self-reported on a benchmark whose full task set is not public and which publishes no confidence intervals. FamilyBench's golden set, red-team set, and both scorecards are downloadable below — raw bytes, hash-verified live — for independent inspection.
- Reproduce it: npm run bench (keyless enforcement suites) and npm run bench:live (armed baseline) in the platform repo — exact commands and downloadable fixtures on the reproducibility page.
The reproducibility bundle
Download the exact golden set, red-team set, and both scorecards — raw JSON, SHA-256 hashes computed live from disk, and a walkthrough for checking the answer key yourself.
Get the bundle →FamilyBench™ is part of FamilyLawGPS by LegalDraft Technologies LLC. Legal information, not legal advice. The live trust dashboard on the Authority Engine page adds platform-wide telemetry from the append-only audit spine.