Get started
Five minutes, one API key, and everything you own that speaks OpenAI — SDKs, opencode, Cursor, your own scripts — gets task-aware routing, sticky conversations, tool calling, metrics and a dashboard. Without changing a line of client code.
The recipe below is deliberately small: two rungs on OpenRouter, both with native tool calling and a very large context. Together they cost a rounding error, so you can leave routing on all day.
1 · install
go install github.com/muthuishere/routsi/cmd/routsi@latest # or: npm install -g @muthuishere/routsi (prebuilt binary, no Go toolchain)
2 · configure
Save as ~/.config/routsi/models.yaml. routsi finds it automatically — a ./models.yaml in the working directory wins, and -config beats both.
${VAR}, including URLs, paths, ports, model names, commands, booleans, integers, and durations. Missing variables fail startup by name; commented references stay inert. Fields ending in _env already contain a variable name and therefore remain unexpanded.listen: ":11080"
default: small
tiers: # what plain `model: "auto"` resolves to
cheap: small
strong: max
models:
- name: small # $0.010 / $0.030 per Mtok · 262k ctx
type: forward
provider: openrouter
base_url: https://openrouter.ai/api/v1
api_key_env: OPENROUTER_API_KEY
upstream_model: inclusionai/ling-2.6-flash
- name: big # $0.030 / $0.130 per Mtok · 1M ctx
type: forward
provider: openrouter
base_url: https://openrouter.ai/api/v1
api_key_env: OPENROUTER_API_KEY
upstream_model: qwen/qwen3.7-flash
- name: max # a NAME, not yet a bigger model
type: forward
provider: openrouter
base_url: https://openrouter.ai/api/v1
api_key_env: OPENROUTER_API_KEY
upstream_model: qwen/qwen3.7-flash
- name: dyn
type: dynamic
levels:
low: small
medium: big
high: maxmax deliberately points at the same upstream as big. A rung is a name, not a model — declaring the top of the ladder up front means you can repoint it later (a bigger model, a local agent, a pull-worker) without touching a single client. Everything keeps asking for dyn.
3 · run
echo 'OPENROUTER_API_KEY=sk-or-...' >> ~/.config/routsi/.env # or just export it routsi serve # routsi on :11080 — 4 models, default small — dashboard http://localhost:11080/
Then point anything at http://localhost:11080/v1 with any non-empty API key.
curl localhost:11080/v1/chat/completions \
-H 'content-type: application/json' \
-d '{"model":"dyn","messages":[{"role":"user","content":"say hi in 3 words"}]}'Make it permanent — routsi install registers a keep-alive service (launchd on macOS, systemd --user on Linux; restart-on-failure, start-at-login, no root).
One name, the right model
Ask for dyn and routsi scores the task and dispatches. The choice is always disclosed in the X-Selected-Model response header, so it is never a black box:
"say hi in 3 words" → X-Selected-Model: small "tradeoffs of optimistic vs pessimistic locking…" → X-Selected-Model: big "…lock-free MPMC ring buffer… prove absence of ABA" → X-Selected-Model: max
Ask for small, big or max by name and routing is bypassed entirely — a concrete model name is always a hard bypass.
medium prefers low first. Leaving medium: out quietly sends medium-scored work to the cheap rung.Tool calling, on every rung
Every rung emits native OpenAI tool_calls — including parallel ones in a single turn — and accepts role: "tool" results back:
finish_reason: tool_calls
call_9f59a57e6481492 get_weather {"city": "Chennai"}
call_c85bb134fcb643a get_weather {"city": "Tokyo"}That matters if you drive routsi from an agent loop: the point is that the tool contract survives the routing. Every routsi transport — forward, translated, pull-worker and exec adapter — is covered by an end-to-end capability matrix in the repo (task matrix).
Sticky conversations
Send an X-Conversation-Id header (or a conversation_id body field) and the conversation pins to whichever rung it landed on. Pins escalate only — a conversation that reached max never silently drops back to small mid-thread. An explicit conversation id also turns on proxy-managed memory: routsi keeps the transcript, so the client sends only the new message each turn.
Growing the config
Everything here is additive — drop it into the same models: list.
Headroom you only reach on purpose. Leave a model out of every group and it can only be reached by name, so it can never surprise the bill. Point dyn.high at it the day the ladder needs real headroom — a one-line change, and every client keeps asking for dyn.
Your own agent session as a model. A pull-worker queue is a routable model like any other — no inbound URL, no port to open. See Pull-workers.
A local script as a model. Anything that reads a job on stdin and writes an answer on stdout is a model — see Agents as models.
Your own routing brain. A decider: block swaps the built-in scorer for any executable — see How auto decides.
Gotchas worth knowing up front
Reasoning models eat max_tokens before they write a word. Ask for three words with max_tokens: 300 and you can get content: null, finish_reason: "length", and 300 reasoning tokens. Two fixes, both verified through routsi:
{"max_tokens": 2000} // give it room
{"reasoning": {"enabled": false}} // or turn thinking offBecause type: forward is raw passthrough, provider-specific knobs like reasoning reach the upstream untouched — routsi rewrites only the model name and the auth header.
Any API key works. routsi is open by default — the client key is discarded and replaced with the upstream key named by api_key_env. To require a token, see Auth & mTLS.
Check the price before you pin a model. Hosted catalogs move; the ids and prices above were correct when written, not forever.