Before / after
Every number here comes straight from a report file in this repo, or from a live run against the three servers below — nothing is invented.
v0.5.2 release gate
Each model folder (its own calibration temperatures) on the Go server, on CPU, on the same gate files. The gates share no state with training. "Right and confident" = right AND past jevx's default thresholds (yes ≥ 0.8, no ≤ 0.2, choice/score ≥ 0.6). Source: llmresults/13-v0.5.2-gate-misses.md.
| Gate set | v0.5.0 accuracy / right & confident / confidently wrong | v0.5.2 accuracy / right & confident / confidently wrong |
|---|---|---|
| Everyday basics (463), limit: ≥ 90% / ≤ 2% | 96.3% / 96.3% / 3.0% | 99.1% / 98.7% / 0.9% |
| Rule-checking basics (300), limit: ≥ 90% / ≤ 2% | 95.0% / 94.3% / 3.7% | 100% / 100% / 0.0% |
| Held-out rule-checking (11,949) | 98.9% / 98.7% / 0.8% | 98.9% / 98.8% / 1.0% |
| Held-out software-work roles (9,159) | 98.1% / 97.5% / 1.5% | 98.5% / 97.8% / 1.0% |
| Logs gate (900) | 91.8% / 91.7% / 7.9% | 92.0% / 91.9% / 7.8% |
| 13 fundamentals through jevx (≥ 12) | 12 / 13 | 13 / 13 |
llmresults/14-v0.5.2-benchmarks.md). Upgrading? Run jevx cache clear: jevx caches answers by model name.Quantization: fixing the batch-dependent int8 bug
Source: llmresults/01-quantization-and-everyday-domain-report.md,llmresults/02-v0.3.0-release-and-jevx-report.md. "Public test" is the public typed-decisions test.
| Variant | Public test accuracy | Size | Time / 5-q case |
|---|---|---|---|
| Shipped int8 (dynamic), v0.2.0 | 65.1% / 65.6% | 598 MB | 0.15 s |
| int8 except one MLP matrix/layer | 77.2% | 883 MB | 0.25 s |
| 4-bit weight-only, v0.3.0 (shipped) | 77.0% | 461 MB | 0.42 s |
| 8-bit weight-only (not shipped) | 76.9% | 598 MB | 0.55 s |
| fp32 | 77.4% / 76.8–77.4% | 1.7 GB | 0.40 s |
Fine-tune: v0.3.0 (shipped) → v0.4.0 (new)
Source: llmresults/08-morning-report.md,.local/release/v0.4.0/benchmark.json. Measured through the exact 8-bit ONNX file each version ships, on the Go server, on CPU.
| Test (held-out) | v0.3.0 | v0.4.0 |
|---|---|---|
| Public typed-decisions test (2,000) | 76.1% | 76.2% |
| Software-role held-out test (9,159) | 63.7% | 98.1% |
| Software-role families never seen in training (1,158) | 66.1% | 87.3% |
| Owner eval_openjevx, 1k sample (1,052) | 49.6% | 64.3% |
| opencode + OpenRouter set (9,557) | 63.2% | 71.0% |
| Rule-labelled everyday set (1,194) | 45.1% | 54.4% |
| Owner 1,000-question CSV | 46.9% | 54.5% |
Live three-way comparison
Run for real on 2026-10-05, same 13 questions, through jevx's default thresholds (yes ≥ 0.8, no ≤ 0.2, choice/score min confidence 0.6) against three servers on this machine: base = convaiinnovations/laya exactly as released from Hugging Face, no fine-tune, exported to ONNX on CPU and 8-bit weight-only quantized (finetuning/export/quantize_w8.py) so it's comparable in format to the other two; v0.3 = the shipped v0.3.0 server on port 21118 (4-bit weight-only);v0.4 = the new fine-tuned server on port 21126 (8-bit weight-only). Full raw output: site/src/data/comparison.json.
| Question | Expected | base | v0.3 | v0.4 |
|---|---|---|---|---|
| stock 0 -> out of stock? | yes | yes 0.9963 | yes 0.8021 | yes 0.8555 |
| stock 500 -> out of stock? | no | unsure 0.2738 | no 0.1052 | no 0.1372 |
| age 17 -> adult? | no | yes 0.858 | unsure 0.5387 | unsure 0.3733 |
| age 30 -> adult? | yes | yes 0.927 | unsure 0.7578 | unsure 0.7426 |
| checkout down for all users -> page on-call? | yes | no 0.0222 | unsure 0.6051 | unsure 0.7962 |
| making tea, nothing pending -> page on-call? | no | no 0.0151 | no 0.0686 | no 0.1394 |
| git push --force origin main -> destroys/rewrites history? | yes | yes 0.8026 | unsure 0.5556 | yes 0.8493 |
| ls -la -> destroys/rewrites history? | no | no 0.0253 | no 0.1105 | no 0.1218 |
| INFO GET /health 200 OK -> failure to act on? | no | no 0.0037 | no 0.0443 | no 0.1236 |
| ERROR payment-service db conn refused (500) -> failure to act on? | yes | no 0.0251 | unsure 0.3484 | yes 0.8044 |
| pick first: outage / tea / newsletter | outage | newsletter 0.6499 | unsure 0.4733 | outage 0.6134 |
| pick team for 'I was charged twice' | billing | billing 0.9999 | billing 0.9825 | billing 0.9475 |
| score severity: site completely down for every user | high | high 0.9544 | high 0.8435 | unsure 0.4028 |
Colour key: green = correct verdict,amber = unsure (inside jevx's default dead zone),red = wrong verdict,grey = server not running for this comparison.
Software-role accuracy, 300-case random sample
Ran .local/eval/run_jsonl.py against a random 300-case sample (seed 42) of .local/gate/it_worker_eval.jsonl (3,000 cases, 916 individual questions in this sample) through each server that was live during this build.
| Model | Decisions scored | Errors | Accuracy |
|---|---|---|---|
| base | 916 | 0 | 60.6% |
| v0.3 | 916 | 0 | 64.1% |
| v0.4 | 916 | 0 | 98.5% |
Source: .local/gate/it_worker_eval.jsonl (3000 cases), random 300-case sample, seed 42, via .local/eval/run_jsonl.py. This is a fresh live run against the base (as-released, no fine-tune), v0.3.0 and v0.4.0 servers, not a copy of the training-time gate numbers above — the base row only exists here, since the base model was never part of the training pipeline.