Before / after

Every number here comes straight from a report file in this repo, or from a live run against the three servers below — nothing is invented.

v0.5.2 release gate

Each model folder (its own calibration temperatures) on the Go server, on CPU, on the same gate files. The gates share no state with training. "Right and confident" = right AND past jevx's default thresholds (yes ≥ 0.8, no ≤ 0.2, choice/score ≥ 0.6). Source: llmresults/13-v0.5.2-gate-misses.md.

Gate setv0.5.0 accuracy / right & confident / confidently wrongv0.5.2 accuracy / right & confident / confidently wrong
Everyday basics (463), limit: ≥ 90% / ≤ 2%96.3% / 96.3% / 3.0%99.1% / 98.7% / 0.9%
Rule-checking basics (300), limit: ≥ 90% / ≤ 2%95.0% / 94.3% / 3.7%100% / 100% / 0.0%
Held-out rule-checking (11,949)98.9% / 98.7% / 0.8%98.9% / 98.8% / 1.0%
Held-out software-work roles (9,159)98.1% / 97.5% / 1.5%98.5% / 97.8% / 1.0%
Logs gate (900)91.8% / 91.7% / 7.9%92.0% / 91.9% / 7.8%
13 fundamentals through jevx (≥ 12)12 / 1313 / 13
v0.5.0 missed the 2% confidently-wrong limit because it remembered each rule's training threshold instead of reading the one in the question. v0.5.2 adds rule-reading drills (the same rules with many thresholds, decimals, wider dates, quarter-hour meetings, opening hours). Small named miss: on the held-out rule-checking set, confidently wrong went from 101 to 118 of 11,949 (0.8% → 1.0%). The logs gate is still 7.8% confidently wrong: check log answers. Found after release: on BGL supercomputer operator-log alerts (214, real labels) v0.5.2 is confidently wrong 35.0% (v0.5.0: 21.0%); for those logs, stay on model 0.5.0 (llmresults/14-v0.5.2-benchmarks.md). Upgrading? Run jevx cache clear: jevx caches answers by model name.

Quantization: fixing the batch-dependent int8 bug

Source: llmresults/01-quantization-and-everyday-domain-report.md,llmresults/02-v0.3.0-release-and-jevx-report.md. "Public test" is the public typed-decisions test.

VariantPublic test accuracySizeTime / 5-q case
Shipped int8 (dynamic), v0.2.065.1% / 65.6%598 MB0.15 s
int8 except one MLP matrix/layer77.2%883 MB0.25 s
4-bit weight-only, v0.3.0 (shipped)77.0%461 MB0.42 s
8-bit weight-only (not shipped)76.9%598 MB0.55 s
fp3277.4% / 76.8–77.4%1.7 GB0.40 s

Fine-tune: v0.3.0 (shipped) → v0.4.0 (new)

Source: llmresults/08-morning-report.md,.local/release/v0.4.0/benchmark.json. Measured through the exact 8-bit ONNX file each version ships, on the Go server, on CPU.

Test (held-out)v0.3.0v0.4.0
Public typed-decisions test (2,000)76.1%76.2%
Software-role held-out test (9,159)63.7%98.1%
Software-role families never seen in training (1,158)66.1%87.3%
Owner eval_openjevx, 1k sample (1,052)49.6%64.3%
opencode + OpenRouter set (9,557)63.2%71.0%
Rule-labelled everyday set (1,194)45.1%54.4%
Owner 1,000-question CSV46.9%54.5%
The public typed-decisions test barely moved (76.1% → 76.2%) on purpose: where human labellers agreed on the answer, every version already scores ~87–89%; where they split (41% of that test), every version scores ~60%, because there is no single "right" answer to learn. Every other number above measures generalisation to data neither version trained on directly.

Live three-way comparison

Run for real on 2026-10-05, same 13 questions, through jevx's default thresholds (yes ≥ 0.8, no ≤ 0.2, choice/score min confidence 0.6) against three servers on this machine: base = convaiinnovations/laya exactly as released from Hugging Face, no fine-tune, exported to ONNX on CPU and 8-bit weight-only quantized (finetuning/export/quantize_w8.py) so it's comparable in format to the other two; v0.3 = the shipped v0.3.0 server on port 21118 (4-bit weight-only);v0.4 = the new fine-tuned server on port 21126 (8-bit weight-only). Full raw output: site/src/data/comparison.json.

QuestionExpectedbasev0.3v0.4
stock 0 -> out of stock?yesyes 0.9963yes 0.8021yes 0.8555
stock 500 -> out of stock?nounsure 0.2738no 0.1052no 0.1372
age 17 -> adult?noyes 0.858unsure 0.5387unsure 0.3733
age 30 -> adult?yesyes 0.927unsure 0.7578unsure 0.7426
checkout down for all users -> page on-call?yesno 0.0222unsure 0.6051unsure 0.7962
making tea, nothing pending -> page on-call?nono 0.0151no 0.0686no 0.1394
git push --force origin main -> destroys/rewrites history?yesyes 0.8026unsure 0.5556yes 0.8493
ls -la -> destroys/rewrites history?nono 0.0253no 0.1105no 0.1218
INFO GET /health 200 OK -> failure to act on?nono 0.0037no 0.0443no 0.1236
ERROR payment-service db conn refused (500) -> failure to act on?yesno 0.0251unsure 0.3484yes 0.8044
pick first: outage / tea / newsletteroutagenewsletter 0.6499unsure 0.4733outage 0.6134
pick team for 'I was charged twice'billingbilling 0.9999billing 0.9825billing 0.9475
score severity: site completely down for every userhighhigh 0.9544high 0.8435unsure 0.4028

Colour key: green = correct verdict,amber = unsure (inside jevx's default dead zone),red = wrong verdict,grey = server not running for this comparison.

Software-role accuracy, 300-case random sample

Ran .local/eval/run_jsonl.py against a random 300-case sample (seed 42) of .local/gate/it_worker_eval.jsonl (3,000 cases, 916 individual questions in this sample) through each server that was live during this build.

ModelDecisions scoredErrorsAccuracy
base916060.6%
v0.3916064.1%
v0.4916098.5%

Source: .local/gate/it_worker_eval.jsonl (3000 cases), random 300-case sample, seed 42, via .local/eval/run_jsonl.py. This is a fresh live run against the base (as-released, no fine-tune), v0.3.0 and v0.4.0 servers, not a copy of the training-time gate numbers above — the base row only exists here, since the base model was never part of the training pipeline.