Get started

Five minutes, one API key, and everything you own that speaks OpenAI — SDKs, opencode, Cursor, your own scripts — gets task-aware routing, sticky conversations, tool calling, metrics and a dashboard. Without changing a line of client code.

The recipe below is deliberately small: two rungs on OpenRouter, both with native tool calling and a very large context. Together they cost a rounding error, so you can leave routing on all day.

1 · install

install
go install github.com/muthuishere/routsi/cmd/routsi@latest
# or: npm install -g @muthuishere/routsi   (prebuilt binary, no Go toolchain)

2 · configure

Save as ~/.config/routsi/models.yaml. routsi finds it automatically — a ./models.yaml in the working directory wins, and -config beats both.

Environment expansion everywhere. Every ordinary YAML scalar supports ${VAR}, including URLs, paths, ports, model names, commands, booleans, integers, and durations. Missing variables fail startup by name; commented references stay inert. Fields ending in _env already contain a variable name and therefore remain unexpanded.
~/.config/routsi/models.yaml
listen: ":11080"
default: small

tiers:            # what plain `model: "auto"` resolves to
  cheap: small
  strong: max

models:
  - name: small                     # $0.010 / $0.030 per Mtok · 262k ctx
    type: forward
    provider: openrouter
    base_url: https://openrouter.ai/api/v1
    api_key_env: OPENROUTER_API_KEY
    upstream_model: inclusionai/ling-2.6-flash

  - name: big                       # $0.030 / $0.130 per Mtok · 1M ctx
    type: forward
    provider: openrouter
    base_url: https://openrouter.ai/api/v1
    api_key_env: OPENROUTER_API_KEY
    upstream_model: qwen/qwen3.7-flash

  - name: max                       # a NAME, not yet a bigger model
    type: forward
    provider: openrouter
    base_url: https://openrouter.ai/api/v1
    api_key_env: OPENROUTER_API_KEY
    upstream_model: qwen/qwen3.7-flash

  - name: dyn
    type: dynamic
    levels:
      low: small
      medium: big
      high: max
api_key_env names an environment variable, never a key. routsi reads it at startup; the value never enters the config file, the logs, or the dashboard.

max deliberately points at the same upstream as big. A rung is a name, not a model — declaring the top of the ladder up front means you can repoint it later (a bigger model, a local agent, a pull-worker) without touching a single client. Everything keeps asking for dyn.

3 · run

run
echo 'OPENROUTER_API_KEY=sk-or-...' >> ~/.config/routsi/.env   # or just export it
routsi serve
# routsi on :11080 — 4 models, default small — dashboard http://localhost:11080/

Then point anything at http://localhost:11080/v1 with any non-empty API key.

curl
curl localhost:11080/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{"model":"dyn","messages":[{"role":"user","content":"say hi in 3 words"}]}'

Make it permanent — routsi install registers a keep-alive service (launchd on macOS, systemd --user on Linux; restart-on-failure, start-at-login, no root).

One name, the right model

Ask for dyn and routsi scores the task and dispatches. The choice is always disclosed in the X-Selected-Model response header, so it is never a black box:

measured, live
"say hi in 3 words"                                → X-Selected-Model: small
"tradeoffs of optimistic vs pessimistic locking…"  → X-Selected-Model: big
"…lock-free MPMC ring buffer… prove absence of ABA" → X-Selected-Model: max

Ask for small, big or max by name and routing is bypassed entirely — a concrete model name is always a hard bypass.

Declare all three levels. An omitted level falls back to the nearest declared one, and medium prefers low first. Leaving medium: out quietly sends medium-scored work to the cheap rung.

Tool calling, on every rung

Every rung emits native OpenAI tool_calls — including parallel ones in a single turn — and accepts role: "tool" results back:

response
finish_reason: tool_calls
  call_9f59a57e6481492  get_weather  {"city": "Chennai"}
  call_c85bb134fcb643a  get_weather  {"city": "Tokyo"}

That matters if you drive routsi from an agent loop: the point is that the tool contract survives the routing. Every routsi transport — forward, translated, pull-worker and exec adapter — is covered by an end-to-end capability matrix in the repo (task matrix).

Sticky conversations

Send an X-Conversation-Id header (or a conversation_id body field) and the conversation pins to whichever rung it landed on. Pins escalate only — a conversation that reached max never silently drops back to small mid-thread. An explicit conversation id also turns on proxy-managed memory: routsi keeps the transcript, so the client sends only the new message each turn.

Growing the config

Everything here is additive — drop it into the same models: list.

Headroom you only reach on purpose. Leave a model out of every group and it can only be reached by name, so it can never surprise the bill. Point dyn.high at it the day the ladder needs real headroom — a one-line change, and every client keeps asking for dyn.

Your own agent session as a model. A pull-worker queue is a routable model like any other — no inbound URL, no port to open. See Pull-workers.

A local script as a model. Anything that reads a job on stdin and writes an answer on stdout is a model — see Agents as models.

Your own routing brain. A decider: block swaps the built-in scorer for any executable — see How auto decides.

Gotchas worth knowing up front

Reasoning models eat max_tokens before they write a word. Ask for three words with max_tokens: 300 and you can get content: null, finish_reason: "length", and 300 reasoning tokens. Two fixes, both verified through routsi:

json
{"max_tokens": 2000}                 // give it room
{"reasoning": {"enabled": false}}    // or turn thinking off

Because type: forward is raw passthrough, provider-specific knobs like reasoning reach the upstream untouched — routsi rewrites only the model name and the auth header.

Any API key works. routsi is open by default — the client key is discarded and replaced with the upstream key named by api_key_env. To require a token, see Auth & mTLS.

Check the price before you pin a model. Hosted catalogs move; the ids and prices above were correct when written, not forever.

Next: Core concepts →