The journey
Four releases, what each one fixed. Full numbers on the results page.
v0.1
Laya, fine-tuned on public data
convaiinnovations/laya trained on the public typed-decisions set: 77.4% on its benchmark.
v0.2
Two problems
- int8 lost 12 points (65.1% vs 76.8% fp32): the same question got a different answer depending on what else was in the batch.
- Near chance on everyday software decisions (46.8%), and too unsure to pass jevx's 0.8 bar even on "stock is 0, is it out of stock?".
v0.3
Fix the quantization
4-bit weight-only: activations stay float, so batching no longer changes answers. 77.0% at 461 MB, no retrain.
v0.4
Fine-tune on real software work — $1.80
443,394 decisions across 7 roles and 56 families, after removing 22,576 questions that leaked from the test sets. One rented RTX 4090, under 3 hours.
Software-role accuracy 63.7% → 98.1%; families never seen in training 66.1% → 87.3%.
Why the public test stays at ~76%. On the 41% of questions where its human labellers disagreed, every version scores ~60% — there's no single right answer to learn. Where they agreed, every version already scores ~87–89%.