Health insurers sit on one of the strangest text corpora in existence. Every claim a member generates — every diagnosis code, procedure code, and prescription fill — lands in a sequence, ordered in time, spanning years. It reads like a language: 4019 (hypertension) tends to precede 42731 (atrial fibrillation); an insulin fill follows a diabetes code the way a verb follows its subject. The foundation-model recipe — pretrain on the raw sequences, fine-tune on the tasks you actually care about — seems almost custom-built for this data.

The literature says it works. Med-BERT, CLMBR, and MOTOR all showed that pretrained clinical-sequence models transfer to downstream prediction tasks — at institutional scale, on real data, with institutional compute. I wanted the one-person version of a narrower, more decision-shaped question: for the two problems a payer actually pays for — care-management targeting and payment integrity — when does a pretrained claims encoder beat well-engineered gradient boosting, and when doesn’t it?

So I pretrained a 17M-parameter encoder on 517,390 synthetic Medicare members, fine-tuned it twice, benchmarked it against XGBoost baselines that were built and frozen before the transformer existed, and wrote down everything — including the parts where the fancy method loses. Total GPU spend, including every mistake and one deliberate mid-run kill: $1.14.

What This Post Covers

Everything here is reproducible from the public repo: config-driven, seeded, with a 43-test suite enforcing the leakage rules and a budget ledger in the README.

Part I — The Bet

The corpus nobody fine-tunes on

A payer’s view of a member is not a chart; it’s a claims stream. That stream is lower-resolution than clinical notes but it has three properties a modeler should love: it’s longitudinal, it’s coded in a closed vocabulary, and it exists for every member, not just the sick ones. The bet behind this project is that the stream itself — not just aggregate features derived from it — carries predictive structure, and that a small transformer can learn it once, unsupervised, and amortize it across tasks.

The honest counter-bet: for tabular-ish problems at payer scale, gradient boosting over well-engineered aggregates is brutally hard to beat, and most of the predictive signal (last year’s spend predicts this year’s spend) doesn’t need sequence modeling at all. Both bets get tested here, and I committed to publishing whichever way it went.

Two datasets and a clean-split trick

The pretraining corpus is CMS DE-SynPUF — synthetic Medicare claims for 2008–2010, released as twenty disjoint samples. That disjointness is a gift: I pretrained on samples 3–7 (517,390 members, 41.3M coded events — measured) and reserved samples 1–2 exclusively for downstream evaluation. Zero member overlap between pretraining and evaluation, by construction, enforced by tests rather than promises.

The fraud dataset is the Kaggle Healthcare Provider Fraud corpus: 558k claims across 5,410 labeled providers — and, crucially, it descends from the same DE-SynPUF lineage, so it speaks the same ICD-9 vocabulary. The entire cross-dataset transfer story hinges on that shared vocabulary, so I refused to spend a GPU dollar until it was measured: after building the 28,203-token pretraining vocabulary, 99.9% of the fraud dataset’s diagnosis-code occurrences were covered (measured, gated at milestone M1). Had that number come back under ~60%, the transfer design would have been redesigned before any training.

DE-SynPUF samples 3–7 517,390 members · 41.3M coded events DE-SynPUF samples 1–2 held out — never pretrained on Kaggle provider fraud 5,410 labeled providers · 558k claims claims-fm encoder 6 layers × d320 · 16.7M params masked-code modeling · $0.52 Task A — member risk admission · top-decile cost Task B — provider fraud + label-efficiency grid Hybrid frozen embeddings ⊕ XGBoost pretrain fine-tune frozen [CLS] labels + eval only provider claim sequences shared ICD-9 vocabulary 99.9% dx coverage — measured gate

One encoder, two tasks, three datasets' worth of discipline. Samples 1–2 never touch pretraining; the Kaggle fraud data connects only through the shared ICD-9 vocabulary — a bridge whose load-bearing capacity was measured (99.9% coverage) before anything was built on it.

Two data traps worth immortalizing. First, CMS’s own portal serves sample 1’s 2010 beneficiary file under a sample 20 filename — the canonical URL 404s, and the mislink has existed since at least 2014. I recovered the true file through Wayback Machine content digests and verified it by member-ID nesting (112,754 IDs, 100% contained in sample 1’s earlier years — measured). Second, the Kaggle files encode missing diagnosis codes as the literal string NA, which my first vocabulary-overlap run happily counted as the single most frequent “diagnosis” in the dataset — 70% of all occurrences. The tell was a statistical impossibility: 90% of code types covered but only 30% of occurrences, exactly backwards from how frequency distributions work. Checksums and inverted-statistics paranoia are not optional in this line of work.

Part II — Building It, One Honest Measurement at a Time

Baselines first, frozen first

Before any transformer code existed, I built the models the transformer would have to beat: logistic regression and XGBoost (40-draw random search) over 159 engineered member features and 34 engineered provider features — cost history, utilization counts, chronic-condition flags, coding-pattern statistics. They were tuned on train/validation only, scored once on test, and frozen under a git tag (v0.2-baselines). Every transformer comparison in this post is against those frozen numbers; nothing got quietly re-tuned after the fact.

The baselines immediately produced the project’s first humbling result: on provider fraud with full labels, plain logistic regression beat tuned XGBoost on test (AUPRC 0.749 vs 0.711 — measured, overlapping CIs). Thirty-four well-chosen features over 5,410 providers are nearly linearly separable; the extra capacity bought variance, not signal. That result stayed in the report, because “the simplest model that wins, wins” is a deployment principle and not an embarrassment.

$0.52 of pretraining

The encoder is deliberately small: 6 layers, d=320, 16.7M parameters, with summed embeddings for code, claim type, age-at-event, and calendar month, and [VISIT] tokens marking visit boundaries. The objective is masked-code modeling — BERT’s recipe with medical codes as the vocabulary. Twelve epochs over 71.6M token positions on a $0.25/hr RTX 4090 spot instance: validation loss fell monotonically (7.52 → 6.70 — measured) and early stopping never fired.

What does a model learn from synthetic claims? More than the frequency table, less than clinical language:

0.1% 1% 10% 50% masked-code top-1 accuracy (log scale) Pretraining: model vs always-guess-the-most-common-code code kind 1.4% 15.8% overall 4.5% 21.6% dx 3.4% 26.4% px 0.14% 10.4% rx frequency prior (modal code) claims-fm encoder (val)

Masked-code top-1 accuracy on held-out members vs the frequency prior (always guessing the most common code of that kind). The encoder is 5–8× the prior on clinical codes and 74× on drug codes — but far below the 40–60% real-EHR models reach, a first hint of the synthetic ceiling. Log scale.

The embedding space passes the sniff test in the way that matters for a payer audience — anchor a code, list its nearest neighbors by cosine similarity:

Anchor Nearest neighbors (cosine)
5859 chronic kidney disease NOS kidney disorder NOS · CKD stage III · CKD stage IV · ESRD · renal failure NOS
25000 type-2 diabetes hypothyroidism · hypercholesterolemia · hyperlipidemia · diabetes-uncontrolled · diabetic neuropathy
V5867 long-term insulin use the entire V58.6x long-term-medication family, in order
3995 hemodialysis venous catheterization · cath for renal dialysis · thoracentesis

The CKD anchor retrieving its own staging ladder — stage III, stage IV, end-stage — is the kind of structure you hope pretraining finds. Nobody told the model those codes are ordinal; co-occurrence within member histories did (measured; probe tables in the repo).

Two engineering potholes are worth their paragraph. On my Mac’s MPS backend, training silently stalled at ~6% CPU: the loss used boolean-mask indexing, which produces a different tensor shape every step, and Apple’s graph compiler re-benchmarks kernels per shape. The fix — fixed-shape cross-entropy with an ignore index — is the standard formulation for a reason, and it helps CUDA kernel caching too. Then the first 4090 run OOM’d a 24GB card at a nominal 16k-token batch: my batch-cost model used unpadded lengths while the collator padded to multiples of 64, letting batches of short sequences overshoot the budget by up to 8×. Both bugs were caught by cheap local smoke tests before they could burn real money; both fixes are commits, not footnotes.

The training itself ran on spot instances that the runbook assumes will die. Checkpoints carry full state — model, optimizer, scheduler, batch cursor, RNG — and rsync to my laptop every two minutes, so the instance is always disposable. I didn’t take that on faith: mid-run, I kill -9‘d the trainer on purpose and resumed from the last checkpoint. The loss curve continued without a spike, because the resume path replays the identical batch and mask streams from the cursor (measured: resumed at step 7,000, re-converged through the same trajectory; there’s also a CPU test asserting the property bitwise). The drill found a real bug the tests couldn’t — checkpoint loading moved the RNG state to the GPU, where PyTorch refuses to restore it — which is exactly why you drill on the real hardware once before trusting the design. The instance’s SSH proxy also died twice during the run; the monitoring fell back to the provider’s API and the checkpoint stream never gapped. Fault tolerance you haven’t exercised is decor.

Member risk: the GBM wins, and the hybrid says why

Task A is care-management targeting: from two years of observation (2008–2009), predict next-year inpatient admission (11.6% prevalence) and top-decile total cost (10.0% by construction) for a 114,041-member cohort. Three transformer arms — a frozen-encoder linear probe, a full fine-tune, and a from-scratch control at identical architecture and budget — against the frozen baselines:

Member next-year risk: AUROC with 95% CIs (test, scored once) 0.65 0.70 0.75 AUROC XGBoost (engineered features) Hybrid: XGB + embeddings Logistic regression Pretrained, full fine-tune Pretrained, frozen probe Transformer, from scratch inpatient admission (prev 11.6%) top-decile cost (prev 10.0%)

AUROC with 95% bootstrap CIs, both Task A labels, test set scored once. XGBoost wins clearly; the pretrained arms beat the from-scratch control; the hybrid lands exactly on the XGBoost baseline — that last fact is the informative one.

Two results live in that chart. The one that stings: tuned XGBoost beats every transformer variant — 0.710 vs 0.669 AUROC on admissions, 0.762 vs 0.715 on cost (measured, non-overlapping CIs). We would ship the GBM. The one that redeems: pretraining transferred — the full fine-tune beats its from-scratch twin on both labels (AUPRC 0.184 vs 0.170, 0.246 vs 0.217 — measured), and even the frozen probe matches or beats from-scratch full training.

But “the transformer lost” and “the sequence channel is empty” are different claims, and the difference matters for anyone deciding whether to try this on real data. So I ran the cheapest decisive experiment I know: re-tune XGBoost with the identical protocol on the engineered features plus the frozen encoder’s 320-dimensional member embedding. Same model family, same search, same seed — only the feature set changes. Result: statistically identical to features-alone (0.763 vs 0.762 AUROC on cost, 0.713 vs 0.710 on admissions — measured). On synthetic member histories, the encoder holds no predictive signal the aggregates don’t already carry. That is a measured ceiling, and it indicts the data, not the method: DE-SynPUF’s synthesis preserves the aggregate statistics engineered features consume while destroying the sequential structure a transformer needs — visible upstream in that 15.8% masked accuracy. On real claims, where the Med-BERT/CLMBR/MOTOR results live, the same $0 experiment is the first thing I’d rerun.

The capture curve is the payer-language version of the same story — if care management can only outreach 5% of members, the frozen XGBoost baseline catches 20.5% of true top-decile-cost members, the transformer 17.0% (measured):

0% 20% 40% 60% share of true high-cost members captured 0 10% 20% 30% share of members outreached, ranked by predicted risk Top-decile cost: capture at outreach capacity (test) random outreach XGBoost pretrained transformer

Capture at outreach capacity, top-decile cost label. The gap between the curves is real members a capacity-limited program would miss — the operational cost of the AUROC difference.

And one result I’d show any team shipping risk scores: the best admission model came out of tuning with severe class-imbalance weighting, which wrecked its probabilities — ECE 0.32, predicted risks wildly inflated relative to observed rates. Isotonic recalibration fit on validation fixed it completely (ECE 0.005 — measured):

0.00 0.25 0.50 0.75 1.00 observed admission rate 0 0.5 1 mean predicted probability Raw model output 0.00 0.25 0.50 0.75 1.00 0 0.5 1 mean predicted probability After isotonic (fit on val)

The same XGBoost admission model, before and after isotonic recalibration (fit on validation only). Discrimination is untouched; the probabilities go from fiction to usable. Payers act on probabilities — outreach lists are sized by them — so this chart is not optional hygiene.

Provider fraud: transfer pays where labels are scarce

Task B is payment integrity: flag potentially fraudulent providers from their claim patterns. Here the encoder crosses datasets — pretrained on DE-SynPUF, fine-tuned on Kaggle providers it has never seen, connected only by that measured vocabulary bridge. A provider becomes a sequence of its claims (each claim a [VISIT]-delimited span of its diagnosis and procedure codes), and the encoder’s [CLS] state becomes the provider representation.

One honesty note before the chart: those sequences truncate at 512 tokens, which keeps only 47% of claims and blinds the transformer to volume — information the baseline’s aggregate features get for free. The pretrained-vs-scratch comparison shares the handicap and stands; the transformer-vs-XGBoost comparison is not information-equal, and I say so rather than pretend otherwise.

The operational reality of fraud detection is that confirmed labels are the scarcest resource in the building — each one is a slow, expensive SIU investigation. So the experiment that matters is not “who wins with all the labels” but “who wins at the label budget you actually have”:

0.55 0.60 0.65 0.70 test AUPRC (fraud, prevalence 9.4%) 10% 25% 100% share of labeled providers used for training Provider fraud: what pretraining buys when labels are scarce Hybrid: XGB + frozen embeddings Transformer, pretrained Transformer, from scratch XGBoost, engineered features

The money chart. Test AUPRC vs share of labeled providers used for training; faint dots are individual subsample seeds. Below full labels the ordering is exactly the pretraining hypothesis: pretrained > from-scratch > XGBoost. At 100%, XGBoost recovers the lead and the hybrid tops the board.

Reading it off (measured, three seeds per point): at 10% of labels, the pretrained encoder leads XGBoost 0.623 to 0.594; at 25%, 0.679 to 0.637 — and the pretrained encoder at a quarter of the labels is approaching what XGBoost needs the full label set to reach (0.711). Stability is half the result: across label subsamples the pretrained arm varies by ±0.004–0.018 AUPRC while XGBoost swings ±0.050–0.056. When labels are scarce, pretraining doesn’t just raise the mean — it makes the outcome repeatable enough to plan around.

The hybrid closes the loop: bolting the frozen provider embeddings onto XGBoost’s features lifts it to 0.688 at 25% labels and 0.718 at 100% — with the best precision@50 of any model in the project, 82% (measured). Review the top fifty providers on that ranking and four in five are true flags. At full labels the honest scoreboard still has simple models on top (that logistic regression again, 0.749), so the deployment answer is a ladder, not a winner: simple models when labels are plentiful, the fine-tuned encoder when they’re scarce and stability matters, embeddings-into-GBM when you want peak triage precision without changing your serving stack.

At an operating point a fraud unit would recognize — review the top 100 providers per cycle — the pretrained transformer delivers 56% precision at 74% recall (measured), meaning an investigator’s queue is majority-true and catches three-quarters of flagged-provider fraud. Two caveats belong next to that number, stated rather than buried. The labels are the dataset author’s constructed flags, not adjudicated SIU outcomes; constructed labels are cleaner and partly rule-like, so every model here — baseline and transformer alike — overstates what confirmed-fraud labels would show. And 66% of beneficiaries appear under more than one provider, so while the splits are clean at the provider level (the unit of prediction), member-level information is not fully disjoint across train and test. Quantified, disclosed, and shared by all models being compared.

The framing is no accident. I spent years in security operations, and provider billing anomalies are anomalous event streams — the same detection posture as security telemetry: baseline the population, rank deviations, respect analyst capacity, and assume the adversary adapts. It’s why the operating metric here is precision at an SIU caseload rather than a threshold-free curve, and why the output is framed as triage support, not automated adjudication.

Process lessons from a $1.14 pipeline

The budget wasn’t the constraint; it was the instrument. Cheap compute forced an ordering — measure locally, gate before spending, verify before trusting — that turned out to be the quality system.

The near-miss that justifies the paranoia. My first Task B evaluation compared calibrated transformer probabilities against the baselines’ raw ones. Isotonic regression fit on an 812-provider validation set creates ranking ties, which quietly depressed the transformer’s AUPRC by 0.06 — in the direction that made my own model look worse, which is how I know the protocol was honest and also how subtle these slips are. The frozen baseline protocol was uncalibrated for ranking metrics; the comparison had to be raw-vs-raw. Caught before freezing, fixed in one commit, and now memorialized in a check: the repo’s verify_blog_numbers.py re-derives every number in this post from the committed metrics artifacts, so the prose can’t drift from the JSONs even if I do.

Gates beat retrospectives. The two moments that most shaped this project were both pre-spend gates: the vocabulary-overlap measurement (which licensed the transfer design before a GPU dollar) and the frozen-baseline tag (which fixed the target before the transformer could negotiate with it). Neither cost anything. Both made later results harder to argue with — including by me.

Spot-instance ops is a discipline, not a vibe. Everything on the rented GPUs assumed eviction: two-minute checkpoint pulls, API-based liveness monitoring for when the SSH proxy flaked (it did, twice), and a deliberate kill/resume drill exercised on the long pretraining run. Total overhead for that discipline across the whole project: pennies. The alternative — re-running a silently-dead job — costs more the first time it happens.

The Evidence Ledger

Every load-bearing number above, tied to its committed artifact in the repo:

Claim Value Source
Pretraining corpus 517,390 members / 41.3M events DATA.md, pretrain_pack/meta.json
Vocabulary bridge 99.9% dx-occurrence coverage reports/vocab_overlap.md
Masked top-1 vs prior 15.8% vs 1.4% reports/pretrain.md, metrics.jsonl
Task A, XGB vs transformer AUROC 0.710/0.762 vs 0.669/0.715 baselines/metrics_task_a.json, reports/metrics_task_a_transformer.json
Hybrid = baseline on Task A 0.713/0.763 AUROC reports/metrics_hybrid.json
Calibration repair ECE 0.32 → 0.005 baselines/metrics_task_a.json
Fraud @ 25% labels 0.679 (±0.004) vs 0.637 (±0.050) reports/metrics_task_b_transformer.json
Fraud @ 100%, hybrid P@50 0.718 AUPRC, 82% reports/metrics_hybrid.json
Total GPU spend $1.14 README budget ledger

A verify_blog_numbers.py script in the repo re-checks this post’s figures against those artifacts; the five charts above are generated from the same JSONs by analysis/blog_svg.py, so the prose and the pixels share a source of truth.

What I Learned

Baselines-first is a forcing function, not a formality. Freezing the GBM under a tag before writing transformer code changed how every later decision felt — there was never a moment where the comparison could drift in the transformer’s favor, which is exactly why the negative results are believable.

Measure the ceiling instead of arguing about it. The hybrid ablation cost nothing and converted “the transformer lost” into “the sequence channel of this dataset is empty, here’s the proof, here’s the cheap test to rerun on data where it isn’t.” Negative results with mechanisms attached are worth more than wins without them.

Label efficiency is the right frame for pretraining. Not “does the transformer beat XGBoost” but “what does pretraining substitute for” — and the answer is labels. That reframe is what makes a claims foundation model interesting to a payer whose bottleneck is confirmed fraud cases, not model capacity.

Synthetic data is a methodology sandbox, not a benchmark. DE-SynPUF let me publish every artifact of this pipeline with zero PHI risk — the right first move — but its synthesis caps what sequence models can show. Absolute numbers here are floors, not estimates.

What I’d Build Next

With real claims: an ICD-10 vocabulary (GEMs-warm-started, validated by the same embedding probes), a next-visit objective à la MOTOR where real temporal structure can reward it, claim-level attention pooling to lift the 512-token provider truncation, and the deployment loop — monthly batch scoring with per-cycle recalibration and drift monitors on population mix, code mix, and calibration decay. And on day one, before any of that: the hybrid test, because it answers in an afternoon whether the sequence channel of your data is worth a transformer at all.


The full build — spec, decision log, leakage tests, budget ledger, and every figure’s source data — is at github.com/nicholicaron/claims-fm.