Worthify research / Decision-1

The base model still matters after you add a LoRA

An early BANKING77 adapter gain reversed in a fresh run with a 388-update budget. The longer result changes what this research preview supports.

Single-seed research preview · Results depend on the task

At 09:00, an operator asks whether the fictional Harbor gateway received a security fix before a noon audit. A vendor advisory says version 5.4.2 contains the fix, but says nothing about Harbor's installation. Decision-1 chooses insufficient evidence from four routes. Later, an installation log and asset inventory arrive. Asked the same question, it chooses supported.

Fictional six-stage replay. Four routes matched the authored expectations; two did not. The model received text state and supplied routes, not live systems or a document search.

The application supplies a text state, a question, and 2–16 described options. The model scores those options in one forward pass. A high option score is not a calibrated probability that a claim is true; the application still decides when to ask for review.

Worthify continued training all 11.9 billion text parameters of Gemma 4 12B on routing, evidence, procedure, uncertainty, and supplied-menu tasks. This is full-weight post-training, followed by separate task LoRAs. One full-weight training seed has completed. Its checkpoint was selected on validation before a sealed 3,200-row test. The planned second full-weight seed remains unfinished, so the results below describe the seed-42 preview.

What happens when another team adapts the model?

We trained fresh, matched rank-16 LoRAs on two bases: original Gemma and Decision-1. Each pair used the same task rows, offered options, prompt, optimizer, update budget, and initial LoRA tensors. Validation selected checkpoints before scoring the official test.

The longer BANKING77 run changes the result. At a 388-update budget, the validation-selected original-Gemma adapter scored 93.54% accuracy on 3,080 official test utterances converted into deterministic 16-option routing menus; the Decision-1 adapter scored 93.21%. The Decision-1 difference was −0.32 percentage points. The second matched seed pair also favored original Gemma: 93.77% versus 92.66%, a −1.10-point difference. Resampling the 77 source intents gave a descriptive 95% interval of −0.94 to +0.29 points for the selected pair, spanning zero. Excluding 28 rows flagged for lexical overlap with full-weight training, validation, or test material left a −0.36-point difference on 3,052 rows. The task-specific improvement gate failed.

This was an exploratory extension chosen after seeing the earlier result. It trained four fresh adapters because optimizer and random-state checkpoints from the 194-update study were unavailable. The earlier, shorter run had favored the Decision-1 adapter: 91.62% versus 90.42%, a +1.20-point difference. Its snapshot is retained below for context; the new run is not a byte-exact continuation. The extended method and results and aggregate report document the protocol and paired outcomes.

BANKING77 validation accuracy over 388 updates for two matched seeds per base. Decision-1 begins lower and catches up late; original Gemma has higher absolute area over the full grid.
Extended 388-update validation grid. Thin lines show each seed; bold lines show two-seed means. Decision-1 has higher baseline-adjusted and late area, while original Gemma has higher absolute area over updates 0–388. The sealed test favored original Gemma. Open full-size chart
Earlier 194-update BANKING77 snapshot: 91.62% with Decision-1 versus 90.42% with original Gemma. CTU-13 macro-F1: 0.115 with Decision-1 versus 0.432 with original Gemma.
Earlier 194-update BANKING77 snapshot, alongside the separate CTU-13 comparison. The longer 388-update BANKING77 result above reversed this routing difference. Tasks and metrics differ. Open full-size chart

The separate CTU-13 comparison also favored original Gemma. In that sealed network-flow scenario, its adapter scored 0.432 macro-F1; the Decision-1 adapter scored 0.115. Both seeds favored original Gemma. BANKING77 is close to intent routing already present in Decision-1's training mix, including CLINC150. The two tasks do not establish a general adapter advantage for Decision-1.

The validation curves need care as well. At 388 updates, Decision-1 had higher baseline-adjusted and late validation accuracy area, while original Gemma had higher absolute area across the full grid from update 0. Both bases already exceeded the protocol's 80% threshold at update 0. Neither run establishes faster convergence. The shorter curve below shows its own 194-update window, where Decision-1 started lower, gained more relative to its baseline, and finished slightly ahead; original Gemma still had higher absolute area.

Earlier 194-update BANKING77 validation accuracy. Thin lines show each of two seeds per base and bold lines show their means. Original Gemma has higher absolute area under the curve; the Decision-1 mean finishes higher.
Earlier 194-update validation grid, retained as an exploratory snapshot. Thin lines show each of two seeds per base; bold lines are seed means. Checkpoint selection used validation data. Open full-size chart

One model, several kinds of choices

The full-weight preview exceeded frozen Gemma on the primary held-out metric in seven evaluated families. In routing macro-F1, it scored 67.29% versus 62.21%; in WANLI-derived evidence, 74.41% versus 67.24%. An earlier task LoRA remained ahead of the full-weight model on WANLI and MNLI evidence point estimates. Four families use authored or supplied-menu fixtures. They test the interface, but cannot establish real-world operational performance.

Seven held-out family primary scores comparing frozen Gemma, earlier task LoRA, and the Decision-1 full-weight seed-42 preview. Decision-1 exceeds frozen Gemma on all seven; the task LoRA leads on two evidence families.
Seven held-out families scored with the same BF16 batch-four path on one A100. Family metrics differ; the chart should be read within each family. Open full-size chart

The checkpoint also makes concrete decisions in small demonstrations. In synthetic traffic windows, it chooses pattern labels such as beacon_like from connection metadata; four of six routes matched authored expectations. Those labels do not identify an attacker. In a falling-block game, code supplies legal placements and board features as text while gravity advances. In one run, the model selected goals and buttons for 30 pieces, cleared six lines, and applied 25 rotations. That run has no matched raw-Gemma control. Neither demonstration sends images or packet payloads to the model.

One recorded, single-seed falling-block run. Code handled legal moves and game physics; Decision-1 selected from supplied text choices.

Try the question on your own task

A useful starting point is a decision your software already makes: describe the current state, enumerate permissible choices, and keep an explicit review route where the consequence requires one. The Decision-1 model card records the model revision, results, and limits. Follow the pinned LoRA training recipe to train and verify an adapter on your own choices, then judge it against a matched original-Gemma control.