The 3-Second Visual Story
1
The Problem
β
2
The Roadblock
β
3
The Engineering Pivot
Six Things That Make This Project Stand Out
Emergency Brakes on Bad Data
The Paperwork Lag Rule
Digital Tamper Seals
AI Advises, Humans Decide
The Placebo (Sugar Pill) Test
Strict Digital Blueprints
Real-World Conservation Operations & Mathematical Formulation
The Real-World Outreach Ladder
How carriers escalate missed payments during the statutory grace period
Days 1β5
1. Billing Hiccup
Days 10β15
2. Grace Notice
Days 20β28
3. Human Phone Call
Day 30+
4. Lapse & Recovery
Formula 1: The Bounded Sigmoid Link (ADR 0012)
λtotal(t) = 0.10 · σ(zlapse) + 0.05 · σ(zsurrender) ≤ 0.1500 < 0.2000
Formula 2: Calibrated Triage Scoring (Phase 2.08)
p̂ = σ(0.961849 · z - 0.033420) | Top 5% Lift: 2.31×
Frequently Asked Questions
Do life insurance customer service teams really call customers about missed payments in real life?
Yes, absolutely. In life insurance, this is an essential operational function called Policy Conservation. Unlike a streaming service, losing life insurance is often irreversibleβreinstating coverage requires medical questionnaires and doctor visits, and older policyholders may become medically uninsurable. Furthermore, insurance agents have an urgent financial incentive to call: if a policy lapses in the first 1 to 2 years, the carrier claws back their upfront commission (known as a commission chargeback). For unassigned "orphan" accounts, carrier call centers use triage queues like Inforsight's Top 5% priority list to phone at-risk policyholders during the 30-day grace period to resolve banking changes before coverage is lost.
In simple words, what is 'Policy Conservation'?
When someone buys life insurance, it costs the insurance company hundreds of dollars to set up the policy. If the customer cancels after 1 year, the company loses money and the customer loses protection. "Conservation" means reaching out to struggling or confused customers before their policy lapses to offer help, like payment grace periods or plan adjustments.
Why not just download real customer data from the internet?
Real insurance records contain sensitive personal details (names, health conditions, social security numbers) which cannot be publicly shared. Furthermore, public datasets are usually just single spreadsheets with no historical timestamps, making it impossible to test whether an AI accidentally looks into the future. Writing our own high-fidelity simulator was the cleanest, most legally sound way to do it.
Why does the project proudly show 52% accuracy instead of 90%+?
Because 52% was the scientific truth about our v1 and v3 data generators! The simulator had too much random noise, so no AI could honestly beat a coin flip. In bad engineering, people tweak settings until the score looks good. In great engineering, you diagnose the root cause (ADR 0006/0007), fix the underlying equations, and keep an unalterable audit trail.
Project Progression: Day-0 to Present
Roadblocks, Mistakes & How We Fixed Them
Scientific Rigor & Root Cause Analysis
The 5 Design Iterations: What We Tried, What Failed, Why, and The Decisions
Architecture Decision Record (ADR) Index
Formal technical agreements that cannot be changed without written peer review.
| ADR | Title | Status | In Plain English |
|---|---|---|---|
0001 |
Clean-Room Synthetic Data | Accepted | No proprietary company secrets or customer personal data allowed. Built 100% from public math & original code. |
0002 |
Separate Risk from Action | Accepted | AI only predicts danger; business rules check customer eligibility; human case officers have final veto. |
0003 |
Start Local, Defer Cloud | Accepted | Prove algorithms work on a regular laptop first. Don't waste money spinning up expensive cloud servers prematurely. |
0004 |
v2 Predeclared Gate | Accepted | Write down test pass/fail rules BEFORE looking at the model's scores to prevent moving the goalposts. |
0005 |
Dual-Time v3 Substrate | Accepted | Block retroactive paperwork leaks: events must be both formally effective AND entered in the system before cutoff. |
0006 |
v4 Diagnostic Boundary | Accepted | Formulate 6 rigorous scientific hypotheses to test why v3 failed before writing new code. |
0007 |
Approve v4 Redesign | Accepted | Double the behavioral signal strength, cut random noise in half, and add realistic scheduled payment records. |
0008 |
Authorize Post-v4 Diagnostics | Accepted | Freeze 17 diagnostics and 320-cell feasibility surface to test if proportional hazards can ever work. |
0009 |
Record R2-14B Readiness Stop | Accepted | Halted fail-closed before result access because Contract 1.0.0 lacked mechanical H1-H5 disposition thresholds. |
0010 |
Amend Diagnostic Contract | Accepted | Froze quantitative numerical truth tables into Contract 1.1.0 to eliminate analyst discretion during execution. |
0011 |
Record v5 Design Infeasibility Stop | Accepted | Proved that 0/320 feasibility cells work under exponential link; halted proportional hazards implementation track permanently under the Trilemma. |
0012 |
Bounded Sigmoid Hazard Link | Accepted | Replaced runaway exponential link with bounded logistic link λ = λmax·σ(z), capping monthly hazard ≤ 0.1500 < 0.2000 and qualifying Generation v6. |
0013 |
Amend v6 Statistical Acceptance Protocol | Accepted | Adopted Protocol 3.1.0 with calibrated secondary rule tolerances (joint binomial coverage, quadrature ε, variance contraction floor), achieving 10/10 rule family pass and mechanical PROCEED. |
Why These Exact Thresholds?
Failed in v3 (0/20 Passed)
1. Minimum Accuracy (AUC β₯ 0.65)
Target: β₯ 0.65 | Observed: 0.519 (Median)
Why 0.65 was chosen:
What actually happened in v3:
Failed in v3 (0/20 Passed)
2. Beating the Placebo (Lift β₯ +0.10)
Target: β₯ +0.10 | Observed: +0.016 (Median)
Why +0.10 over Null was chosen:
What actually happened in v3:
Failed in v3 (0/20 Passed)
3. Statistical Consistency (16 of 20 Seeds Must Pass)
Target: β₯ 16 of 20 (80%) | Observed: 0 of 20 (0%)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
π΄ 0 Seeds Passed
βͺ 20 Seeds Required β₯ 16 Passing (80%)
Why 16 of 20 (80%) was chosen:
What actually happened in v3:
Passed in v3 (146β173 per fold)
4. Adequate Event Counts (β₯ 50 Positives per Fold)
Target: β₯ 50 Cases | Observed: 146 to 173 Cases
Why 50 cases was chosen:
What actually happened in v3:
Passed in Phase 2.08 (Platt Scaling)
5. Calibration Quality (ECE ≤ 0.0300 & Slope ∈ [0.85, 1.15])
Target: ECE ≤ 0.0300 | Slope ∈ [0.85, 1.15] | Observed: ECE 0.0115 | Slope 0.9498
Why ≤ 0.0300 ECE was chosen:
What actually happened in Phase 2.08:
Metrics & Validation Observatory
Phase 1 / 2 (v1 Substrate)
100 Policies
Observations: 100 first-billing
Logistic Accuracy: 0.564 (AUC)
XGBoost Accuracy: 0.533 (AUC)
Decision: Pipeline Engineering Only
Phase 2R (v2 Substrate)
3,600 Policies
Observations: 42,795 recurring
Logistic Accuracy: 0.543 (AUC)
XGBoost Accuracy: 0.551 (AUC)
R2-07 Gate: STOP (Leakage Detected)
Phase 2R (v3 Substrate)
14,400 Policies
Observations: 76,545 dual-time
Acceptance Seeds: 20 signal/null pairs
Median Signal: 0.519 (vs Null 0.507)
R2-11 Gate: REDESIGN (Weak Signal)
Phase 2R (v6 Candidate Selected)
0.7057 AUC
Selected Architecture: Logistic Regression
Comparator Model: XGBoost (0.6801 AUC)
Calibration (Brier): 0.1287 (Logistic) vs 0.1332
Point-in-Time Features: 17 features (0 leak flags)
R2-15 Gate: CANDIDATE FROZEN (Logistic)
Phase 2R.16A (v6 Acceptance Gate)
0.7031 AUC
Tested Seeds: 20 reserved pairs (120 units)
Signal Consistency: 20 / 20 pass (AUC β₯ 0.65)
AP Lift / Brier Skill: +0.1344 / +0.0658
Mechanical Decision: PROCEED (10/10 Rule Families)
R2-16A Gate: MECHANICAL PROCEED
Phase 2.08 (Probability Calibration)
0.0115 ECE
Selected Calibrator: Platt Scaling (A=0.962, B=-0.033)
Discrimination: 0.6998 AUC (Exact Rank Preservation)
Calibration Slope: 0.9498 (Within [0.85, 1.15])
Brier Score: 0.1211 (vs Uncalibrated 0.1212)
Top 1% Review Queue: 34.09% Precision (2.23x Lift, NNR 2.9)
Top 5% Review Queue: 11.57% Recall (2.31x Lift, 35.31% Prec)
P2-08 Gate: OPERATIONAL READY (4 Risk Tiers)
Phase 2.11 (Final Evaluation & Release)
RELEASE
Out-of-Sample AUC: 0.6998 [0.6847, 0.7153] (Gate G1)
Average Precision: 0.2765 [0.2560, 0.2994] (Gate G2)
Calibration ECE / Slope: 0.0115 (G3) / 0.9498 (G4)
Top 1% Precision: 34.09% (G5 >= 0.3000, 2.23x lift)
Top 5% Lift / Recall: 2.31x (G6 >= 2.00x) / 11.57%
Gates G1βG6: 6 of 6 Passed (100%)
P2-11 Decision: OFFICIAL RELEASE
Empirical Discrimination: Model vs. Placebo Control
Placebo / Matched Null
Logistic Baseline
XGBoost
Target Threshold (0.65)
Engineering Rigor & Test Suite Health
395
Automated Safety Tests
12
Contract Schema Tests
100%
Boundary Verification
0
Test Set Peeking Incidents
System Architecture & The Authority Model
The 4-Tier Authority Hierarchy
The Paperwork Delay Rule
WHERE effective_at <= as_of
AND ingested_at <= as_of
Clean-Room Isolation
./scripts/check_repository_boundaries.sh
# Status: Passed (0 violations)
Interactive Pipeline Simulator
1
Boundary Audit
Zero Secrets & IP
2
Synthetic Gen
Fictional Events
3
Dual-Time Filter
Point-in-Time Lock
4
Auth Scoring
Tamper Digest Seal
5
Model Fitting
Train on Past Only
6
Placebo Gate
Matched Null Test
7
Final Decision
Proceed / Stop
Stage 1 of 7
Ready to Begin
Select a scenario above and click "Run Pipeline Simulation" to watch the animated process step-by-step.
What happens here:
bash β Inforsight Automated Test Harness
$ make check
Awaiting execution trigger...