- Rust 92.4%
- Nix 7.6%
| assets | ||
| nix | ||
| src | ||
| .gitignore | ||
| AGENTS.md | ||
| Cargo.lock | ||
| Cargo.toml | ||
| flake.lock | ||
| flake.nix | ||
| README.md | ||
oc — free opencode.ai models over OpenAI API
No account needed. Accepts requests like LM Studio / any OpenAI client,
impersonates opencode and talks to https://opencode.ai/zen/v1.
Run
./target/release/oc
# or
PORT=1234 HOST=127.0.0.1 ./target/release/oc
Environment variables:
| var | default | meaning |
|---|---|---|
PORT |
1234 |
port (same as LM Studio) |
HOST |
127.0.0.1 |
bind address |
ZEN_BASE |
https://opencode.ai/zen/v1 |
upstream |
OPENCODE_API_KEY |
public |
key; public without an account. Secrets better via OPENCODE_API_KEY_FILE (file wins) |
PROXIES |
— | outbound proxy pool: socks5h://user:pass@host:port, http://... separated by ,/;/newline |
PROXIES_FILE |
— | file with proxy list (one URL per line, # comments), merged with PROXIES |
PROXY_FAILOVER_AFTER |
3 |
consecutive retryable failures before switching to proxies (round-robin). No errors — always direct |
PROXY_COOLDOWN_SECS |
180 |
how long to stay on proxies once failed over (proxy successes extend the window, direct success ends it) |
MAX_UPSTREAM_ATTEMPTS |
3 |
sends per logical turn, skipping cooling-down routes with backoff; only transport errors and 429/5xx are retried |
RETRY_DELAY_MS |
800 |
base delay between upstream attempts (exponential: base, 2x, 4x, ...) |
RATE_LIMIT_COOLDOWN_SECS |
120 |
base per-route block after a 429, doubled per repeat (Retry-After wins when present) |
RATE_LIMIT_MAX_SECS |
43200 |
cap for the grown 429 block |
ROUTE_ERROR_COOLDOWN_SECS |
30 |
per-route block after a transport error or 5xx |
PROBE_MODELS |
all free | static probe models; live client models from the last 24h are added automatically |
USE_DIRECT |
1 |
include a direct outbound (0 = proxies only) |
EGRESS_ORDER |
sticky |
sticky prefers direct, spread round-robins healthy routes |
MIN_OUTPUT_TOKENS |
1024 |
floor for client output caps, tiny caps always return incomplete (0 = off) |
MODELS_CACHE_SECS |
300 |
upstream /v1/models cache TTL (0 = off) |
DEFAULT_MODEL |
nemotron-3-ultra-free |
model when the client doesn't specify one |
MAX_HISTORY_CHARS |
auto | history cap in chars; unset = computed per model (see below) |
MAX_HISTORY_MESSAGES |
500 |
history cap in messages (0 = unlimited) |
MAX_TOOL_ROUNDS |
3 |
tool-loop follow-ups per request (0 = single attempt, still sanitized) |
TOOL_UNAVAILABLE_MSG |
— | custom text fed back as every tool result (default: no workspace access, answer directly) |
FORWARD_CLIENT_TOOLS=1 |
0 |
append client tools upstream; pure client-tool rounds are returned to the client for execution |
SHOW_ALL_MODELS=1 |
— | also list paid models in /v1/models |
LOG_LEVEL=basic |
off |
trace verbosity: off/0, basic/1, full/2, trace/3 (DEBUG_MODE, DEBUG are aliases) |
LOG_MAX_CHARS=2000 |
2000 |
preview cap per dump (0 = unlimited; trace is always unlimited; DEBUG_MAX_CHARS is an alias) |
RUST_LOG=oc=debug |
— | verbose logs (session hit/miss visible) |
Log levels
For "why did it answer strangely" investigations — every client request,
upstream request, turn result and tool call is logged at info level
(visible without touching RUST_LOG):
LOG_LEVEL=basic— one line per client request (model, stream, session, message roles/sizes), one line per upstream turn (route,ses_/msg_, latency, model/sections/caps), one line per turn result (finish reason, SSE bytes, text preview,usage, tool names+args), plus history truncation notices. System prompt is shown as[system Nch], never dumped.LOG_LEVEL=full— everything frombasic, plus full client JSON, upstream JSON (system redacted to length) and raw upstream SSE, all capped byLOG_MAX_CHARS(dumps usemax*10).LOG_LEVEL=trace— wire mode: full untruncated JSON and SSE, including the ~20KB system prompt on every turn.
LOG_LEVEL=basic ./target/release/oc
LOG_LEVEL=full LOG_MAX_CHARS=5000 ./target/release/oc
LOG_LEVEL=trace ./target/release/oc 2>&1 | tee strange-dialog.log
NixOS: services.oc-zen-proxy.logLevel = "full"; (+ logMaxChars).
Secrets are never logged (no Authorization header, no proxy passwords).
Connect LM Studio
Base URL: http://127.0.0.1:1234/v1, models are discovered automatically
(muse-spark-*-contributor-free, mimo-v2.5-free,
ling-3.0-flash-fin-free, nemotron-3-ultra-free,
nemotron-3.5-lightning-free).
deepseek-v4-flash-free is not advertised — down on the zen side
at the time of writing (400/500 on every format).
To pin a dialog to a fixed upstream session, send a custom header
x-opencode-session: ses_... (LM Studio supports custom headers).
Check:
curl http://127.0.0.1:1234/v1/models
curl http://127.0.0.1:1234/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"nemotron-3-ultra-free","messages":[{"role":"user","content":"What is 2+2?"}]}'
Which URL goes where
The proxy serves the whole namespace on one port (PORT, default 1234):
| what | where to put it | example |
|---|---|---|
| Base URL for OpenAI clients (LM Studio, Cline, Continue, Open WebUI, aichat) | Base URL / API Base field | http://127.0.0.1:1234/v1 |
Full chat URL (tgpt --url, curl) |
— | http://127.0.0.1:1234/v1/chat/completions |
| Model list | — | http://127.0.0.1:1234/v1/models |
| Behind nginx on a server | proxyPass |
http://127.0.0.1:1234/ (root to root, clients append /v1/... themselves) |
Note: Base URL always ends with /v1, while nginx proxyPass has
no suffix (root to root). API key is public everywhere (or empty if
the client allows it). Via server: https://your-domain/v1/chat/completions.
How the mimicry works
The free tier (Bearer public) answers 403 FreeTierError to anyone
not looking like opencode. The proxy signs every request as opencode run:
Authorization: Bearer public,User-Agent: opencode/...x-opencode-session: ses_.../x-opencode-request: msg_...— valid IDs perpackages/schema/src/identifier.ts(first 12 hex =timestamp_ms * 0x1000 + counter, then 14 base62); random IDs are rejected by timestamp check.msg_is fresh per request,ses_is sticky per dialog (see below)stream: trueupstream always (withstream:falseupstream answers403, so for non-streaming clients SSE is reassembled back into onechat.completion)- the body is wrapped in the exact
opencode runtemplate:assets/main_system.txt(~20KB system) +assets/tools.json(11 built-in tools,tool_choice:auto) + client messages. Missing template tools or a foreign system —403(verified live); extra tools appended to the template are tolerated (also verified live). Client tools are appended upstream only withFORWARD_CLIENT_TOOLS=1, otherwise ignored. Upstreamtool_callsreach the client only in forward mode, see "Tool-call loop" below.
Assets were captured by intercepting a real opencode run
(OPENCODE_MODELS_PATH -> local logger).
Per-model format routing
Zen serves different models in different wire formats (wrong format =
generic 500 Internal server error — that's what aichat hit on muse-spark).
The proxy routes itself, the client always speaks OpenAI chat:
| models | upstream | translation |
|---|---|---|
nemotron-*-free, mimo-*, ling-* |
POST /zen/v1/chat/completions |
SSE as is |
muse-spark-*-contributor-free |
POST /zen/v1/responses |
Responses SSE → chat SSE (text, reasoning, function calls, usage) |
Empty client API key is fine, the proxy ignores it
(upstream always gets Bearer public or OPENCODE_API_KEY).
History caps (dynamic)
No worst-case-for-all limits. On startup the proxy fetches
models.opencode.ai/api.json (same catalog opencode uses) and takes each
model's limit.context (unreachable → minimal 200k assumed for all,
refresh on restart). Per request:
history = context − template − response_reserve − 10% margin
where template is measured from our own assets (~11k tokens),
response_reserve is the request's max_tokens (or 32k default).
Result in chars (×4), floored at 20k. So muse-spark (1M window) gets
~900k tokens of history while mimo (200k) gets ~150k. Explicit
MAX_HISTORY_CHARS wins over the computed value.
Trimming keeps the newest messages (oldest dropped first) and never
leaves a dangling head (a tool result without its call goes too).
It runs before the template is prepended, so system+tools stay intact.
Second layer is reactive: on overflow-ish upstream errors (400/413
mentioning context/length) the proxy halves the cap and retries in the
same session (prefix cache stays warm), up to 2 retries.
Sticky sessions (cache)
Upstream keeps a prefix cache (system+tools) and sticky provider routing
keyed by x-opencode-session, like real opencode. The OpenAI protocol
has no sessions, so the proxy derives them: hash of (model + history) —
a retry hits the full hash, the next dialog turn hits the history hash
minus the last 1-3 messages (clients usually append
[assistant reply + user]). Forks/edits/new chats transparently get
a fresh session. Cache: 512 dialogs, 30 min TTL. Verified in logs:
the second turn reuses the same ses_ (session hit, depth=2).
Tool-call loop
Upstream answers with opencode tool calls (read, bash, …) that the
client never asked for and cannot run. Forwarding them as-is kills plain
chats (empty reply / error), so the proxy loops instead: every call is
answered with "tool unavailable, answer directly" and the request is
resent (same session, fresh msg_), until clean text arrives.
usage is summed across rounds. If the model keeps calling tools past
MAX_TOOL_ROUNDS, the client gets the accumulated text, or a short note
if there is none — opencode tool_calls never leak to the client.
MAX_TOOL_ROUNDS=0 ./target/release/oc # single attempt, no follow-ups
Forwarding client tools (FORWARD_CLIENT_TOOLS=1)
Off by default. When on, client-declared tools are appended upstream
next to the template (verified: extras don't 403, verified on both
chat/completions and responses), so the model can actually call them.
A round that calls only client tools is returned to the client as-is
(tool_calls + finish_reason: tool_calls, streamed as deltas) — the
client executes them and sends back role=tool results, the loop
continues upstream. Calls to template tools (or unknown names, or mixed
rounds) are still stubbed as above, but with forwarding on the stub then
names the client tools the model may use instead (with forwarding off the
classic message is kept, so the model isn't tempted to hallucinate calls
to tools it cannot see). In NixOS:
services.oc-zen-proxy.forwardClientTools = true;.
Proxy failover + retries
Every outbound path (direct + each proxy) has its own cooldown: a 429
parks that route for RATE_LIMIT_COOLDOWN_SECS (default 120s, Retry-After
wins when present — upstream FreeUsageLimitError carries no timing itself),
a transport error or 5xx parks it for ROUTE_ERROR_COOLDOWN_SECS
(default 30s). Each logical turn is sent up to MAX_UPSTREAM_ATTEMPTS
times (default 3) with exponential backoff (RETRY_DELAY_MS, default 800ms),
skipping parked routes — so a rate-limited direct is followed by proxy
attempts within the same client request, and a route that failed minutes
ago is not retried until its cooldown expires. When every route is parked,
the turn does one live parallel recheck of all egress and uses whoever
answers; only if nobody does, the client immediately gets the last real
upstream 429 body verbatim, with no further waiting.
Retried: transport errors and 429/500/502/503/504. Permanent errors
(400/401/403/404/422, incl. geo-block 403) return immediately.
A counter of consecutive retryable failures at PROXY_FAILOVER_AFTER
switches the preferred order to proxies-first for PROXY_COOLDOWN_SECS
(default 180s): proxy successes keep the window open, the first direct
success resets the counter and returns to direct. Switches are visible
in logs (route rate-limited, cooling down / failing over to proxies /
retrying upstream request ... / back to direct, passwords never logged).
/v1/models always goes direct (it has a static fallback).
Egress health
Every outbound path (direct + each proxy) cools down on failures: a 429
parks that route (RATE_LIMIT_COOLDOWN_SECS, doubled per repeat up to
RATE_LIMIT_MAX_SECS), a transport error or 5xx parks it briefly
(ROUTE_ERROR_COOLDOWN_SECS). Limits are per model and IP, so probes run
one inference request (template + tiny cap) per PROBE_MODELS model across
both wire formats, and a route counts as healthy only if every probed model
answers; /models alone is never probed, it is not rate-limited and reports
healthy on dead IPs. Probes run on manual /v1/update and automatically as
a live recheck when every route is parked; traffic successes unpark at once.
Manual refresh: curl host/v1/update deep-probes every egress in parallel,
one IP MODEL STATUS line each, no credentials:
138.99.37.196 nemotron-3-ultra-free 200. Dead routes park themselves,
recovered ones rejoin rotation at once.
Pool status without probing: curl host/v1/health lists alive egress
(one host per line, parked ones omitted).
export PROXIES="socks5h://u1:p1@1.2.3.4:1080,socks5h://u2:p2@5.6.7.8:1080"
export PROXY_FAILOVER_AFTER=2
./target/release/oc
# or via file: PROXIES_FILE=/run/secrets/oc-proxies
Secrets outside git and the store
proxiesFile/apiKeyFile in the module are plain strings, not nix paths:
the file is read at systemd runtime and never copied anywhere. Keep it
e.g. in ~/secrets/socks5 — no need to commit it, no --impure needed:
services.oc-zen-proxy.proxiesFile = "/home/lily/secrets/socks5";
Format: one URL per line, # comments. Two requirements: the file must
exist on the target machine and be readable by the service user — if home
is closed off (chmod 700), allow traversal:
chmod o+x /home/lily /home/lily/secrets && chmod o+r /home/lily/secrets/socks5
(or ACLs via setfacl). For production, sops-nix/agenix with files
in /run/secrets/ is the right call.
In the NixOS module it's even easier: set proxiesFile and the rest
happens by itself — on every nixos-rebuild switch the dir and file
are created if missing (template inside) and chowned to the service
user. Content of an existing file is never touched.
Uncomment the lines, add your proxies, then systemctl restart oc-zen-proxy.
NixOS
nix build .#default
Wire up the module (flake input oc) — the package is substituted
automatically, no overlay needed:
{
inputs.oc.url = "path:/home/lily/Documents/oc"; # or git URL
# ...
imports = [ oc.nixosModules.default ];
services.oc-zen-proxy = {
enable = true;
port = 1234;
host = "127.0.0.1"; # localhost only, nginx faces outward; 0.0.0.0 listens everywhere
apiKeyFile = "/run/secrets/oc-api-key"; # or apiKey = "..."
proxiesFile = "/var/lib/oc/proxy"; # auto-created with a template if missing
failoverAfter = 3;
};
}
Public access via plain nginx: https://ai.example.com/v1
The module doesn't manage nginx — just drop a standard virtualHost next to it:
services.oc-zen-proxy = {
enable = true;
host = "127.0.0.1"; # localhost only, nginx faces outward
port = 1234;
proxiesFile = "/run/secrets/oc-proxies";
};
services.nginx = {
enable = true;
virtualHosts."ai.example.com" = {
forceSSL = true;
enableACME = true;
locations."/" = {
proxyPass = "http://127.0.0.1:1234/";
proxyWebsockets = true;
recommendedProxySettings = true;
# SSE streaming + slow free-tier turns (tool rounds can take minutes).
extraConfig = ''
proxy_buffering off;
proxy_read_timeout 300s;
proxy_send_timeout 300s;
'';
};
};
};
Check: curl https://ai.example.com/v1/models.
All options are in nix/module.nix.
Limitations
- ~10k prompt tokens of overhead per request — same as opencode itself sends, the proxy adds nothing on top (and even saves the TUI title requests). But unlike opencode, LM Studio doesn't compact history — long dialogs hit the context sooner.
- Model persona is opencode (coding assistant), not neutral.
- Reasoning deltas (
nemotron) are folded intoreasoning_contentin non-stream responses. - LM Studio client tools are not proxied (exact opencode set required). Upstream tool calls are auto-answered as unavailable (see "Tool-call loop"), so coding questions get a text answer instead of workspace actions.
- Each tool round re-sends the growing history: multi-round turns cost extra prompt tokens.
- Free models sometimes answer
503 Upstream overloaded— retry. The free roster rotates, some have geo-blocks. - Template is hardcoded: if opencode updates system/tools/UA, expect
403until assets are re-captured.