I
Inforsight v0.3.1 Hardening β€’ RH-13 PROCEED
See Risk. Shape Action. Honest AI for life insurance conservation.
Wording:
GitHub
Executive Summary & Plain-English Walkthrough

An Honest AI System That Refuses to Cheat

56
Merged Code Increments
100% verified with automated tests
395
Automated Safety Tests
Zero future data leaks tolerated
0
Faked Performance Claims
Proudly embraced "Stop" and "Redesign"
320
Feasibility Cells Tested
0 Feasible → Fail-Closed Stop
πŸš€ Direct Page Jump:

The 3-Second Visual Story

1
⚠️

The Problem

βž”
2
πŸ›‘

The Roadblock

βž”
3
πŸ†

The Engineering Pivot

Six Things That Make This Project Stand Out

πŸ›‘

Emergency Brakes on Bad Data

⏱️

The Paperwork Lag Rule

πŸ”’

Digital Tamper Seals

βš–οΈ

AI Advises, Humans Decide

πŸ§ͺ

The Placebo (Sugar Pill) Test

πŸ“œ

Strict Digital Blueprints

Real-World Conservation Operations & Mathematical Formulation

The Real-World Outreach Ladder How carriers escalate missed payments during the statutory grace period
Days 1–5
1. Billing Hiccup
πŸ’³
Days 10–15
2. Grace Notice
πŸ“¬
Days 20–28
3. Human Phone Call
πŸ“ž
Day 30+
4. Lapse & Recovery
πŸ“‹
πŸ›‘οΈ

Formula 1: The Bounded Sigmoid Link (ADR 0012)

λtotal(t) = 0.10 · σ(zlapse) + 0.05 · σ(zsurrender) ≤ 0.1500 < 0.2000
🎯

Formula 2: Calibrated Triage Scoring (Phase 2.08)

p̂ = σ(0.961849 · z - 0.033420)  |  Top 5% Lift: 2.31×
Read full operational walkthrough, commission chargebacks & mathematical proofs →

Frequently Asked Questions

Do life insurance customer service teams really call customers about missed payments in real life?
Yes, absolutely. In life insurance, this is an essential operational function called Policy Conservation. Unlike a streaming service, losing life insurance is often irreversibleβ€”reinstating coverage requires medical questionnaires and doctor visits, and older policyholders may become medically uninsurable. Furthermore, insurance agents have an urgent financial incentive to call: if a policy lapses in the first 1 to 2 years, the carrier claws back their upfront commission (known as a commission chargeback). For unassigned "orphan" accounts, carrier call centers use triage queues like Inforsight's Top 5% priority list to phone at-risk policyholders during the 30-day grace period to resolve banking changes before coverage is lost.
In simple words, what is 'Policy Conservation'?
When someone buys life insurance, it costs the insurance company hundreds of dollars to set up the policy. If the customer cancels after 1 year, the company loses money and the customer loses protection. "Conservation" means reaching out to struggling or confused customers before their policy lapses to offer help, like payment grace periods or plan adjustments.
Why not just download real customer data from the internet?
Real insurance records contain sensitive personal details (names, health conditions, social security numbers) which cannot be publicly shared. Furthermore, public datasets are usually just single spreadsheets with no historical timestamps, making it impossible to test whether an AI accidentally looks into the future. Writing our own high-fidelity simulator was the cleanest, most legally sound way to do it.
Why does the project proudly show 52% accuracy instead of 90%+?
Because 52% was the scientific truth about our v1 and v3 data generators! The simulator had too much random noise, so no AI could honestly beat a coin flip. In bad engineering, people tweak settings until the score looks good. In great engineering, you diagnose the root cause (ADR 0006/0007), fix the underlying equations, and keep an unalterable audit trail.

Project Progression: Day-0 to Present

Roadblocks, Mistakes & How We Fixed Them

Scientific Rigor & Root Cause Analysis

The 5 Design Iterations: What We Tried, What Failed, Why, and The Decisions

Architecture Decision Record (ADR) Index

Formal technical agreements that cannot be changed without written peer review.

ADR Title Status In Plain English
0001 Clean-Room Synthetic Data Accepted No proprietary company secrets or customer personal data allowed. Built 100% from public math & original code.
0002 Separate Risk from Action Accepted AI only predicts danger; business rules check customer eligibility; human case officers have final veto.
0003 Start Local, Defer Cloud Accepted Prove algorithms work on a regular laptop first. Don't waste money spinning up expensive cloud servers prematurely.
0004 v2 Predeclared Gate Accepted Write down test pass/fail rules BEFORE looking at the model's scores to prevent moving the goalposts.
0005 Dual-Time v3 Substrate Accepted Block retroactive paperwork leaks: events must be both formally effective AND entered in the system before cutoff.
0006 v4 Diagnostic Boundary Accepted Formulate 6 rigorous scientific hypotheses to test why v3 failed before writing new code.
0007 Approve v4 Redesign Accepted Double the behavioral signal strength, cut random noise in half, and add realistic scheduled payment records.
0008 Authorize Post-v4 Diagnostics Accepted Freeze 17 diagnostics and 320-cell feasibility surface to test if proportional hazards can ever work.
0009 Record R2-14B Readiness Stop Accepted Halted fail-closed before result access because Contract 1.0.0 lacked mechanical H1-H5 disposition thresholds.
0010 Amend Diagnostic Contract Accepted Froze quantitative numerical truth tables into Contract 1.1.0 to eliminate analyst discretion during execution.
0011 Record v5 Design Infeasibility Stop Accepted Proved that 0/320 feasibility cells work under exponential link; halted proportional hazards implementation track permanently under the Trilemma.
0012 Bounded Sigmoid Hazard Link Accepted Replaced runaway exponential link with bounded logistic link λ = λmax·σ(z), capping monthly hazard ≤ 0.1500 < 0.2000 and qualifying Generation v6.
0013 Amend v6 Statistical Acceptance Protocol Accepted Adopted Protocol 3.1.0 with calibrated secondary rule tolerances (joint binomial coverage, quadrature ε, variance contraction floor), achieving 10/10 rule family pass and mechanical PROCEED.

Why These Exact Thresholds?

πŸ†
Generation v6 Resolution (Phase 2R.16A β€” Protocol 3.1.0 & ADR 0013)

Failed in v3 (0/20 Passed)
1. Minimum Accuracy (AUC β‰₯ 0.65)
Target: β‰₯ 0.65 | Observed: 0.519 (Median)
Coin Flip (0.50)
Too Weak
Acceptable (0.65+)
Observed: 0.519
Target: 0.65
Why 0.65 was chosen:

What actually happened in v3:

Failed in v3 (0/20 Passed)
2. Beating the Placebo (Lift β‰₯ +0.10)
Target: β‰₯ +0.10 | Observed: +0.016 (Median)
Random Noise (+0.00)
Marginal
Real Signal (+0.10+)
Observed: +0.016
Target: +0.10
Why +0.10 over Null was chosen:

What actually happened in v3:

Failed in v3 (0/20 Passed)
3. Statistical Consistency (16 of 20 Seeds Must Pass)
Target: β‰₯ 16 of 20 (80%) | Observed: 0 of 20 (0%)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
πŸ”΄ 0 Seeds Passed βšͺ 20 Seeds Required β‰₯ 16 Passing (80%)
Why 16 of 20 (80%) was chosen:

What actually happened in v3:

Passed in v3 (146–173 per fold)
4. Adequate Event Counts (β‰₯ 50 Positives per Fold)
Target: β‰₯ 50 Cases | Observed: 146 to 173 Cases
Sample Too Small (< 50)
Statistically Sufficient (50+)
Observed: ~160
Target: 50
Why 50 cases was chosen:

What actually happened in v3:

Passed in Phase 2.08 (Platt Scaling)
5. Calibration Quality (ECE ≤ 0.0300 & Slope ∈ [0.85, 1.15])
Target: ECE ≤ 0.0300 | Slope ∈ [0.85, 1.15] | Observed: ECE 0.0115 | Slope 0.9498
Calibrated (≤ 0.03)
Miscalibrated
Severe Error (> 0.06)
Observed: 0.0115
Target: 0.0300
Why ≤ 0.0300 ECE was chosen:

What actually happened in Phase 2.08:

Metrics & Validation Observatory

Phase 1 / 2 (v1 Substrate)
100 Policies
Observations: 100 first-billing
Logistic Accuracy: 0.564 (AUC)
XGBoost Accuracy: 0.533 (AUC)
Decision: Pipeline Engineering Only

Phase 2R (v2 Substrate)
3,600 Policies
Observations: 42,795 recurring
Logistic Accuracy: 0.543 (AUC)
XGBoost Accuracy: 0.551 (AUC)
R2-07 Gate: STOP (Leakage Detected)

Phase 2R (v3 Substrate)
14,400 Policies
Observations: 76,545 dual-time
Acceptance Seeds: 20 signal/null pairs
Median Signal: 0.519 (vs Null 0.507)
R2-11 Gate: REDESIGN (Weak Signal)

Phase 2R (v6 Candidate Selected)
0.7057 AUC
Selected Architecture: Logistic Regression
Comparator Model: XGBoost (0.6801 AUC)
Calibration (Brier): 0.1287 (Logistic) vs 0.1332
Point-in-Time Features: 17 features (0 leak flags)
R2-15 Gate: CANDIDATE FROZEN (Logistic)

Phase 2R.16A (v6 Acceptance Gate)
0.7031 AUC
Tested Seeds: 20 reserved pairs (120 units)
Signal Consistency: 20 / 20 pass (AUC β‰₯ 0.65)
AP Lift / Brier Skill: +0.1344 / +0.0658
Mechanical Decision: PROCEED (10/10 Rule Families)
R2-16A Gate: MECHANICAL PROCEED

Phase 2.08 (Probability Calibration)
0.0115 ECE
Selected Calibrator: Platt Scaling (A=0.962, B=-0.033)
Discrimination: 0.6998 AUC (Exact Rank Preservation)
Calibration Slope: 0.9498 (Within [0.85, 1.15])
Brier Score: 0.1211 (vs Uncalibrated 0.1212)
Top 1% Review Queue: 34.09% Precision (2.23x Lift, NNR 2.9)
Top 5% Review Queue: 11.57% Recall (2.31x Lift, 35.31% Prec)
P2-08 Gate: OPERATIONAL READY (4 Risk Tiers)

Phase 2.11 (Final Evaluation & Release)
RELEASE
Out-of-Sample AUC: 0.6998 [0.6847, 0.7153] (Gate G1)
Average Precision: 0.2765 [0.2560, 0.2994] (Gate G2)
Calibration ECE / Slope: 0.0115 (G3) / 0.9498 (G4)
Top 1% Precision: 34.09% (G5 >= 0.3000, 2.23x lift)
Top 5% Lift / Recall: 2.31x (G6 >= 2.00x) / 11.57%
Gates G1–G6: 6 of 6 Passed (100%)
P2-11 Decision: OFFICIAL RELEASE

Empirical Discrimination: Model vs. Placebo Control

Placebo / Matched Null Logistic Baseline XGBoost Target Threshold (0.65)
0.70 0.65 (Target) 0.60 0.55 0.50 (Coin Flip) v1 (Phase 2) 100 policies v2 (R2-06) Leakage Stop v3 Acceptance Redesign Gate v6 Acceptance 0.7031 PROCEED

Engineering Rigor & Test Suite Health

395 Automated Safety Tests
12 Contract Schema Tests
100% Boundary Verification
0 Test Set Peeking Incidents

System Architecture & The Authority Model

The 4-Tier Authority Hierarchy

Tier 1: Perception

Predictive Model (AI)

Estimates cancellation risk score

Hard Rule: Strictly gives advice. Cannot contact a customer or change a policy.
βž”
Tier 2: Policy Rules

Deterministic Rule Engine

Checks eligibility & laws

Hard Rule: Cannot make up facts. Checks grace periods, regulatory rules, and policy status.
βž”
Tier 3: Synthesis

AI Evidence Assistant

Assembles readable summary

Hard Rule: Must cite exact dates and payment IDs with footnotes. No hallucinations.
βž”
Tier 4: Authority

Human Case Officer

Licensed decision maker

Sole Authority: Signs off on the intervention. Every decision is cryptographically logged.
⏳

The Paperwork Delay Rule

WHERE effective_at <= as_of
  AND ingested_at <= as_of
πŸ—‚οΈ

Clean-Room Isolation

./scripts/check_repository_boundaries.sh
# Status: Passed (0 violations)

Interactive Pipeline Simulator

1 ⏳
Boundary Audit
Zero Secrets & IP
2 ⏳
Synthetic Gen
Fictional Events
3 ⏳
Dual-Time Filter
Point-in-Time Lock
4 ⏳
Auth Scoring
Tamper Digest Seal
5 ⏳
Model Fitting
Train on Past Only
6 ⏳
Placebo Gate
Matched Null Test
7 ⏳
Final Decision
Proceed / Stop
Stage 1 of 7

Ready to Begin

πŸ§ͺ
Select a scenario above and click "Run Pipeline Simulation" to watch the animated process step-by-step.

What happens here:

bash β€” Inforsight Automated Test Harness
$ make check
Awaiting execution trigger...