Fine-tune it on your decisions
The shipped model does not know your rules: that your week is four days, that a discount starts above Rs 500, that repo billing belongs to the payments team. Write those decisions in a spreadsheet, and one command trains a model that does. Each step below says what you should see.
0. What you get and what it costs
You get your own openjevx.w8.onnx (8-bit, at most 750 MB) plus a trainable checkpoint. The model runs on your machine; your decisions never go to a hosted model.
Training runs on a rented GPU (an RTX 4090 on vast.ai, about $0.40–0.60 an hour). One run costs about $1–3. For reference, the v0.5.2 full run took about 2.5 hours on 791k decisions and cost about $1.25; the smoke run before it costs cents. The box deletes itself when it is done.
1. Install
You need Go, Task, uvand Python 3.11+ on your machine.
git clone https://github.com/muthuishere/openjevx.git
cd openjevx
go version && task --version && uv --version && python3 --versionFor the GPU step you also need:
- the
vastaiCLI and a vast.ai API key with a few dollars of credit (pip install vastai && vastai set api-key YOUR_KEY_HERE); - a Cloudflare account for R2 storage: the GPU box downloads your data from R2 and uploads the model back there.
cd finetuning && task cloud-setupcreates the private bucket and the self-destroy endpoint, then checks it; you should see the endpoint answer.
Optional: jevx (for the 13-fundamentals check in the gate and for using the model afterwards) and Docker. Steps 2–4 need no GPU and no cloud account.
2. What goes in your CSV
A row is one decision you want the model to make.
state= the facts at the moment of deciding: what your code, log or ticket knows.question= the decision, phrased as you would ask a colleague.answer= what the right decision was, from your rule or from what really happened. Never a guess.
Start from the example. It covers a 4-day week, 3 working hours a day, a late parcel, a discount over Rs 500, team routing by repo and ticket urgency:
Download decisions.csv (also in the repo: finetuning/examples/decisions.csv, with a one-page column guide)
The columns
| column | required? | what to put | good example | common mistake |
|---|---|---|---|---|
state | yes | the facts, as a JSON object or plain text; include the threshold if it varies | {"hours_logged_today": 3.1} | just true, or leaving out the fact the rule needs |
question | yes | one decision, as you'd ask a colleague | Is the parcel late? | two decisions in one: "is it late and who handles it" |
type | no (noul) | noul = yes/no, choice = pick one, score = rating | choice | using choice for yes/no |
options | choice, score | the keys the model picks between, |-separated; optional short meaning after =; score levels lowest first | payments=billing repos|web=frontend repos | score levels out of order (high|low|medium) |
answer | yes | yes/no (also true/false, 1/0); one option key; a level name or index (0 = lowest) | yes, payments, high | an answer that is not one of the options |
split | no | train, test or gate; empty = 90% train, 10% gate, fixed per row | gate | no gate rows, so you can't measure the result |
source | no | a tag for where the row came from | shop | — |
JSON in a CSV cell: wrap the cell in double quotes and double the quotes inside ("{""day"": ""Friday""}"). Any spreadsheet does this when you save as CSV.
One example per type: the CSV row, and what the model receives
Each row becomes one request to the server, POST /v1/systemone, with your state and question. Training teaches the model to give your answer to that request.
Yes/no (noul)
state,question,type,options,answer
"{""order_total_rs"": 501}","Does the order get the discount? Orders over Rs 500 get 10% off.",noul,,yes{
"state": {"order_total_rs": 501},
"questions": {"q1": {"type": "noul",
"instructions": "Does the order get the discount? Orders over Rs 500 get 10% off."}}
}Pick one (choice)
state,question,type,options,answer
"{""repo"": ""billing"", ""title"": ""Invoice PDF missing tax line""}",Which team owns this?,choice,payments=billing and payments-core repos|platform=infra-terraform and api-gateway repos|web=web-frontend and admin-console repos,payments{
"state": {"repo": "billing", "title": "Invoice PDF missing tax line"},
"questions": {"q1": {"type": "choice", "instructions": "Which team owns this?",
"criteria": {"payments": "billing and payments-core repos",
"platform": "infra-terraform and api-gateway repos",
"web": "web-frontend and admin-console repos"}}}
}Rating (score)
state,question,type,options,answer
"{""customers_affected"": 40, ""payments_failing"": false, ""workaround"": false}",How urgent is this ticket?,score,low|medium|high|critical,high{
"state": {"customers_affected": 40, "payments_failing": false, "workaround": false},
"questions": {"q1": {"type": "score", "instructions": "How urgent is this ticket?",
"criteria": ["low", "medium", "high", "critical"]}}
}Where the facts come from
Put in the state exactly what your system will send the model when it asks for real.
| source | state |
|---|---|
| a log alert | {"line": "OOMKilled payments-api", "env": "prod", "errors_last_10m": 42} |
| a support ticket | {"text": "Charged twice for order 8812", "customer_tier": "gold"} |
| a database record | {"promised_by": "2026-10-03", "delivered_on": "2026-10-04"} |
| a pull request | {"files_changed": 14, "touches": ["migrations/"], "summary": "adds a column"} |
| a form | {"leave_days_requested": 6, "leave_days_left": 4} |
Do and don't
| do | don't |
|---|---|
| near-miss pairs at every threshold: 499 → no, 500 → no, 501 → yes; Thursday → yes, Friday → no | put the answer inside the state ("eligible": true) |
| balance the answers: no single answer above 80% of a question's rows | train on another model's guesses: labels come from the rule or from reality |
| 20+ rows per question | give the same state and question two different answers (a fact is missing) |
keep a few of the hardest rows as split=gate | copy one wording everywhere; vary how the question and facts are phrased |
How many rows
Start with 200–1,000 rows. Each rule becomes many rows: both sides of the threshold, different values, different wordings. The import step warns you when a question has fewer than 20 rows or one answer takes more than 80% of them. The example CSV is deliberately small (68 rows) to show the shape; it gets those warnings too.
3. Import and check
python3 finetuning/dataprep/import_csv.py yours.csv --name mydata --add-to-configYou should see the count of good and bad rows, then one line per question with its answers:
yours.csv: 68 good rows, 0 bad
question rows labels
Does the order get the discount? Orders over Rs 500 get 10% 12 false=6, true=6
Which team owns this? 13 platform=5, payments=4, web=4
warning: 'Which team owns this?': only 13 rows; aim for 20+ with near-miss pairs around the rule
adapter: every row accepted
wrote 53 train rows -> ~/openjevx/data/train/mydata_train.jsonl
wrote 15 gate rows -> ~/openjevx/data/gate/mydata_gate.jsonl
updated ~/.config/openjevx/config.json (backup: config.json.bak-…)A bad row is reported with its line number and a fix (line 7: answer 'maybe' is not yes/no). Nothing is written until every row is good, unless you pass --skip-bad. Use--dry-run to check without writing.
--add-to-config backs up ~/.config/openjevx/config.json, then adds your train file to the training mix (repeated 3 times; change with --repeat), your gate file to the gate, and both to the leakage check. Re-running it does not add duplicates.
4. Test the current model on your gate first
Before paying for a GPU, see how the shipped model already does on your rules. Download the v0.5.2 model folder, then point the gate at it:
curl -L -O https://github.com/muthuishere/openjevx/releases/download/v0.5.11/openjevx-model-0.5.2.tar.gz
tar -xzf openjevx-model-0.5.2.tar.gz # -> model/ (openjevx.w8.onnx, config.json, tokenizer.json)
cd finetuning
task gate -- ../modelThis builds and serves the model on a free local port and scores every gate file, one line each:
gate/mydata_gate.jsonl n= 15 accuracy 60.0% right&confident 40.0% confidently WRONG 13.3%That line is your baseline. If your rules already score well, you may not need to train at all.
5. Train
cd finetuning
task allIt runs these stages and stops at the first failure:
- dataprep generates the built-in rule-labelled sets (basics, conditions, IT work, logs).
- leakage check removes any training question that also appears in a test or gate file.
- trainer check on CPU builds the shard and runs rows through the real trainer, so a data bug fails here, for free, not on a rented GPU.
- smoke run: about 10 minutes on a GPU to prove the whole path works.
- full run: the real training, a few hours.
- gate on the new model.
The GPU box downloads the shard from your private R2 bucket, trains, exports and quantizes to 8-bit, uploads the model, checkpoint and log back to R2, then destroys itself. A timer destroys it at the deadline no matter what. Your laptop only waits on R2, so it can sleep; results wait in the bucket. You should see RUN_DIR=… and, at the end, the path to your newopenjevx.w8.onnx under ~/openjevx/data/work/runs/.
The box clones a pushed commit, so commit and push any change underfinetuning/ before training. Your data goes through R2, not git.
6. Read the gate report
For each gate file you get three numbers:
- accuracy: the top answer was right.
- right & confident: right and sure enough to act on (yes ≥ 0.8, no ≤ 0.2, a choice or score ≥ 0.6). This is what matters: an unsure answer gets escalated, not acted on.
- confidently wrong: sure and wrong. The dangerous one; keep it near zero.
The model ships only if the everyday basics stay at least 90% right & confident, at most 2% confidently wrong, and it gets 12 of jevx's 13 fundamentals. These thresholds are ingate in ~/.config/openjevx/config.json. The full report is saved to~/openjevx/data/work/gate/<model>.json. Compare your line with the baseline from step 4.
7. Run your model
Put the model in a folder with its settings and tokenizer:
my-model/
openjevx.w8.onnx
config.json (its own calibration temperatures)
tokenizer.jsonPoint the server at the folder in openjevx.json and start it:
{
"listen": "127.0.0.1:21160",
"device": "auto",
"model": "/path/to/my-model"
}With Docker, mount the folder and set the same "model" path inside the container. Then add a jevx profile with a versioned model name, so cached answers from the old model are not reused:
jevx profile add mymodel http://127.0.0.1:21160/v1/systemone --model mymodel-v1
jevx profile use mymodel
jevx cache clear
jevx ask --noul discount="Does the order get the discount? Orders over Rs 500 get 10% off." --in '{"order_total_rs": 501}'You should see a yes with high confidence. Train again later? Bump the name to mymodel-v2.
8. Troubleshooting
- The GPU box is stuck loading. Some hosts take long to start. After 15 minutes the box is destroyed and the next machine is tried, up to 3.
- Your network dropped. The box does not depend on your laptop; results wait in R2. Run the step again to pick them up.
- The trainer rejects a row. The CPU trainer check in step 5 names the row and source before any GPU is rented. Fix it in your CSV and import again.
- The gate fails. Look at which file failed. If it is your file, add near-miss rows and check the state holds every fact; if it is the basics, lower
--repeatso your data does not crowd them out.
Downloads for v0.5.2
| What | File | Use it for |
|---|---|---|
| Model folder (467 MB) | openjevx-model-0.5.2.tar.gz | run it: openjevx.w8.onnx + config.json (its calibration temperatures) + tokenizer.json; gate it (step 4) |
| Trainable checkpoint (777 MB) | openjevx-finetune-v0.5.2.tar.gz(sha256) | fine-tune further: model.safetensors, encoder/, tokenizer/, rl_agent_config.json; no training data |
| Hugging Face | muthuishere/openjevx | the same model folder files and checkpoint; v0.5.0 and v0.4.0 are git tags there |
| Example CSV | decisions.csv | the starting point for step 2 |
The v0.5.2 run: one RTX 4090 on vast.ai, 791,239 decisions, one full pass in 2.5 h at about 87 items/s, about $1.25. Test questions were removed first (22,576 leaked questions), and every row was checked against the trainer before renting.
Questions or a bug in the pipeline: open an issue.