The journey

Four releases, what each one fixed. Full numbers on the results page.

v0.1

Laya, fine-tuned on public data

convaiinnovations/laya trained on the public typed-decisions set: 77.4% on its benchmark.

v0.2

Two problems

  • int8 lost 12 points (65.1% vs 76.8% fp32): the same question got a different answer depending on what else was in the batch.
  • Near chance on everyday software decisions (46.8%), and too unsure to pass jevx's 0.8 bar even on "stock is 0, is it out of stock?".
v0.3

Fix the quantization

4-bit weight-only: activations stay float, so batching no longer changes answers. 77.0% at 461 MB, no retrain.

v0.4

Fine-tune on real software work — $1.80

443,394 decisions across 7 roles and 56 families, after removing 22,576 questions that leaked from the test sets. One rented RTX 4090, under 3 hours.

Software-role accuracy 63.7% → 98.1%; families never seen in training 66.1% → 87.3%.

Why the public test stays at ~76%. On the 41% of questions where its human labellers disagreed, every version scores ~60% — there's no single right answer to learn. Where they agreed, every version already scores ~87–89%.

See all the numbers →