At 09:00, an operator asks whether the fictional Harbor gateway received a security fix before a noon audit. A vendor advisory says version 5.4.2 contains the fix, but says nothing about Harbor's installation. Decision-1 chooses insufficient evidence from four routes. Later, an installation log and asset inventory arrive. Asked the same question, it chooses supported.
The application supplies a text state, a question, and 2–16 described options. The model scores those options in one forward pass. A high option score is not a calibrated probability that a claim is true; the application still decides when to ask for review.
Worthify continued training all 11.9 billion text parameters of Gemma 4 12B on routing, evidence, procedure, uncertainty, and supplied-menu tasks. This is full-weight post-training, followed by separate task LoRAs. One full-weight training seed has completed. Its checkpoint was selected on validation before a sealed 3,200-row test. The planned second full-weight seed remains unfinished, so the results below describe the seed-42 preview.
What happens when another team adapts the model?
We trained fresh, matched rank-16 LoRAs on two bases: original Gemma and Decision-1. Each pair used the same task rows, offered options, prompt, optimizer, update budget, and initial LoRA tensors. Validation selected checkpoints before scoring the official test.
The longer BANKING77 run changes the result. At a 388-update budget, the validation-selected original-Gemma adapter scored 93.54% accuracy on 3,080 official test utterances converted into deterministic 16-option routing menus; the Decision-1 adapter scored 93.21%. The Decision-1 difference was −0.32 percentage points. The second matched seed pair also favored original Gemma: 93.77% versus 92.66%, a −1.10-point difference. Resampling the 77 source intents gave a descriptive 95% interval of −0.94 to +0.29 points for the selected pair, spanning zero. Excluding 28 rows flagged for lexical overlap with full-weight training, validation, or test material left a −0.36-point difference on 3,052 rows. The task-specific improvement gate failed.
This was an exploratory extension chosen after seeing the earlier result. It trained four fresh adapters because optimizer and random-state checkpoints from the 194-update study were unavailable. The earlier, shorter run had favored the Decision-1 adapter: 91.62% versus 90.42%, a +1.20-point difference. Its snapshot is retained below for context; the new run is not a byte-exact continuation. The extended method and results and aggregate report document the protocol and paired outcomes.


The separate CTU-13 comparison also favored original Gemma. In that sealed network-flow scenario, its adapter scored 0.432 macro-F1; the Decision-1 adapter scored 0.115. Both seeds favored original Gemma. BANKING77 is close to intent routing already present in Decision-1's training mix, including CLINC150. The two tasks do not establish a general adapter advantage for Decision-1.
The validation curves need care as well. At 388 updates, Decision-1 had higher baseline-adjusted and late validation accuracy area, while original Gemma had higher absolute area across the full grid from update 0. Both bases already exceeded the protocol's 80% threshold at update 0. Neither run establishes faster convergence. The shorter curve below shows its own 194-update window, where Decision-1 started lower, gained more relative to its baseline, and finished slightly ahead; original Gemma still had higher absolute area.

One model, several kinds of choices
The full-weight preview exceeded frozen Gemma on the primary held-out metric in seven evaluated families. In routing macro-F1, it scored 67.29% versus 62.21%; in WANLI-derived evidence, 74.41% versus 67.24%. An earlier task LoRA remained ahead of the full-weight model on WANLI and MNLI evidence point estimates. Four families use authored or supplied-menu fixtures. They test the interface, but cannot establish real-world operational performance.

The checkpoint also makes concrete decisions in small demonstrations. In synthetic traffic windows, it chooses pattern labels such as beacon_like from connection metadata; four of six routes matched authored expectations. Those labels do not identify an attacker. In a falling-block game, code supplies legal placements and board features as text while gravity advances. In one run, the model selected goals and buttons for 30 pieces, cleared six lines, and applied 25 rotations. That run has no matched raw-Gemma control. Neither demonstration sends images or packet payloads to the model.
Try the question on your own task
A useful starting point is a decision your software already makes: describe the current state, enumerate permissible choices, and keep an explicit review route where the consequence requires one. The Decision-1 model card records the model revision, results, and limits. Follow the pinned LoRA training recipe to train and verify an adapter on your own choices, then judge it against a matched original-Gemma control.