No description
  • Rust 92.4%
  • Nix 7.6%
Find a file
2026-09-20 19:31:44 +03:00
assets first commit 2026-09-18 09:54:36 +03:00
nix forward client tools upstream, passthrough pure client-tool rounds 2026-09-19 16:05:31 +03:00
src retry responses turns cut by transport, error note instead of empty stop 2026-09-20 19:31:44 +03:00
.gitignore first commit 2026-09-18 09:54:36 +03:00
AGENTS.md Retry via proxies, sticky failover, minimal comments 2026-09-18 15:15:44 +03:00
Cargo.lock first commit 2026-09-18 09:54:36 +03:00
Cargo.toml first commit 2026-09-18 09:54:36 +03:00
flake.lock first commit 2026-09-18 09:54:36 +03:00
flake.nix Retry via proxies, sticky failover, minimal comments 2026-09-18 15:15:44 +03:00
README.md list client tools in stub only when forwarding is on 2026-09-19 16:19:18 +03:00

oc — free opencode.ai models over OpenAI API

No account needed. Accepts requests like LM Studio / any OpenAI client, impersonates opencode and talks to https://opencode.ai/zen/v1.

Run

./target/release/oc
# or
PORT=1234 HOST=127.0.0.1 ./target/release/oc

Environment variables:

var default meaning
PORT 1234 port (same as LM Studio)
HOST 127.0.0.1 bind address
ZEN_BASE https://opencode.ai/zen/v1 upstream
OPENCODE_API_KEY public key; public without an account. Secrets better via OPENCODE_API_KEY_FILE (file wins)
PROXIES — outbound proxy pool: socks5h://user:pass@host:port, http://... separated by ,/;/newline
PROXIES_FILE — file with proxy list (one URL per line, # comments), merged with PROXIES
PROXY_FAILOVER_AFTER 3 consecutive retryable failures before switching to proxies (round-robin). No errors — always direct
PROXY_COOLDOWN_SECS 180 how long to stay on proxies once failed over (proxy successes extend the window, direct success ends it)
MAX_UPSTREAM_ATTEMPTS 3 sends per logical turn, skipping cooling-down routes with backoff; only transport errors and 429/5xx are retried
RETRY_DELAY_MS 800 base delay between upstream attempts (exponential: base, 2x, 4x, ...)
RATE_LIMIT_COOLDOWN_SECS 120 base per-route block after a 429, doubled per repeat (Retry-After wins when present)
RATE_LIMIT_MAX_SECS 43200 cap for the grown 429 block
ROUTE_ERROR_COOLDOWN_SECS 30 per-route block after a transport error or 5xx
PROBE_MODELS all free static probe models; live client models from the last 24h are added automatically
USE_DIRECT 1 include a direct outbound (0 = proxies only)
EGRESS_ORDER sticky sticky prefers direct, spread round-robins healthy routes
MIN_OUTPUT_TOKENS 1024 floor for client output caps, tiny caps always return incomplete (0 = off)
MODELS_CACHE_SECS 300 upstream /v1/models cache TTL (0 = off)
DEFAULT_MODEL nemotron-3-ultra-free model when the client doesn't specify one
MAX_HISTORY_CHARS auto history cap in chars; unset = computed per model (see below)
MAX_HISTORY_MESSAGES 500 history cap in messages (0 = unlimited)
MAX_TOOL_ROUNDS 3 tool-loop follow-ups per request (0 = single attempt, still sanitized)
TOOL_UNAVAILABLE_MSG — custom text fed back as every tool result (default: no workspace access, answer directly)
FORWARD_CLIENT_TOOLS=1 0 append client tools upstream; pure client-tool rounds are returned to the client for execution
SHOW_ALL_MODELS=1 — also list paid models in /v1/models
LOG_LEVEL=basic off trace verbosity: off/0, basic/1, full/2, trace/3 (DEBUG_MODE, DEBUG are aliases)
LOG_MAX_CHARS=2000 2000 preview cap per dump (0 = unlimited; trace is always unlimited; DEBUG_MAX_CHARS is an alias)
RUST_LOG=oc=debug — verbose logs (session hit/miss visible)

Log levels

For "why did it answer strangely" investigations — every client request, upstream request, turn result and tool call is logged at info level (visible without touching RUST_LOG):

  • LOG_LEVEL=basic — one line per client request (model, stream, session, message roles/sizes), one line per upstream turn (route, ses_/msg_, latency, model/sections/caps), one line per turn result (finish reason, SSE bytes, text preview, usage, tool names+args), plus history truncation notices. System prompt is shown as [system Nch], never dumped.
  • LOG_LEVEL=full — everything from basic, plus full client JSON, upstream JSON (system redacted to length) and raw upstream SSE, all capped by LOG_MAX_CHARS (dumps use max*10).
  • LOG_LEVEL=trace — wire mode: full untruncated JSON and SSE, including the ~20KB system prompt on every turn.
LOG_LEVEL=basic ./target/release/oc
LOG_LEVEL=full LOG_MAX_CHARS=5000 ./target/release/oc
LOG_LEVEL=trace ./target/release/oc 2>&1 | tee strange-dialog.log

NixOS: services.oc-zen-proxy.logLevel = "full"; (+ logMaxChars).

Secrets are never logged (no Authorization header, no proxy passwords).

Connect LM Studio

Base URL: http://127.0.0.1:1234/v1, models are discovered automatically (muse-spark-*-contributor-free, mimo-v2.5-free, ling-3.0-flash-fin-free, nemotron-3-ultra-free, nemotron-3.5-lightning-free). deepseek-v4-flash-free is not advertised — down on the zen side at the time of writing (400/500 on every format).

To pin a dialog to a fixed upstream session, send a custom header x-opencode-session: ses_... (LM Studio supports custom headers).

Check:

curl http://127.0.0.1:1234/v1/models
curl http://127.0.0.1:1234/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"nemotron-3-ultra-free","messages":[{"role":"user","content":"What is 2+2?"}]}'

Which URL goes where

The proxy serves the whole namespace on one port (PORT, default 1234):

what where to put it example
Base URL for OpenAI clients (LM Studio, Cline, Continue, Open WebUI, aichat) Base URL / API Base field http://127.0.0.1:1234/v1
Full chat URL (tgpt --url, curl) — http://127.0.0.1:1234/v1/chat/completions
Model list — http://127.0.0.1:1234/v1/models
Behind nginx on a server proxyPass http://127.0.0.1:1234/ (root to root, clients append /v1/... themselves)

Note: Base URL always ends with /v1, while nginx proxyPass has no suffix (root to root). API key is public everywhere (or empty if the client allows it). Via server: https://your-domain/v1/chat/completions.

How the mimicry works

The free tier (Bearer public) answers 403 FreeTierError to anyone not looking like opencode. The proxy signs every request as opencode run:

  • Authorization: Bearer public, User-Agent: opencode/...
  • x-opencode-session: ses_... / x-opencode-request: msg_... — valid IDs per packages/schema/src/identifier.ts (first 12 hex = timestamp_ms * 0x1000 + counter, then 14 base62); random IDs are rejected by timestamp check. msg_ is fresh per request, ses_ is sticky per dialog (see below)
  • stream: true upstream always (with stream:false upstream answers 403, so for non-streaming clients SSE is reassembled back into one chat.completion)
  • the body is wrapped in the exact opencode run template: assets/main_system.txt (~20KB system) + assets/tools.json (11 built-in tools, tool_choice:auto) + client messages. Missing template tools or a foreign system — 403 (verified live); extra tools appended to the template are tolerated (also verified live). Client tools are appended upstream only with FORWARD_CLIENT_TOOLS=1, otherwise ignored. Upstream tool_calls reach the client only in forward mode, see "Tool-call loop" below.

Assets were captured by intercepting a real opencode run (OPENCODE_MODELS_PATH -> local logger).

Per-model format routing

Zen serves different models in different wire formats (wrong format = generic 500 Internal server error — that's what aichat hit on muse-spark). The proxy routes itself, the client always speaks OpenAI chat:

models upstream translation
nemotron-*-free, mimo-*, ling-* POST /zen/v1/chat/completions SSE as is
muse-spark-*-contributor-free POST /zen/v1/responses Responses SSE → chat SSE (text, reasoning, function calls, usage)

Empty client API key is fine, the proxy ignores it (upstream always gets Bearer public or OPENCODE_API_KEY).

History caps (dynamic)

No worst-case-for-all limits. On startup the proxy fetches models.opencode.ai/api.json (same catalog opencode uses) and takes each model's limit.context (unreachable → minimal 200k assumed for all, refresh on restart). Per request:

history = context − template − response_reserve − 10% margin

where template is measured from our own assets (~11k tokens), response_reserve is the request's max_tokens (or 32k default). Result in chars (×4), floored at 20k. So muse-spark (1M window) gets ~900k tokens of history while mimo (200k) gets ~150k. Explicit MAX_HISTORY_CHARS wins over the computed value.

Trimming keeps the newest messages (oldest dropped first) and never leaves a dangling head (a tool result without its call goes too). It runs before the template is prepended, so system+tools stay intact.

Second layer is reactive: on overflow-ish upstream errors (400/413 mentioning context/length) the proxy halves the cap and retries in the same session (prefix cache stays warm), up to 2 retries.

Sticky sessions (cache)

Upstream keeps a prefix cache (system+tools) and sticky provider routing keyed by x-opencode-session, like real opencode. The OpenAI protocol has no sessions, so the proxy derives them: hash of (model + history) — a retry hits the full hash, the next dialog turn hits the history hash minus the last 1-3 messages (clients usually append [assistant reply + user]). Forks/edits/new chats transparently get a fresh session. Cache: 512 dialogs, 30 min TTL. Verified in logs: the second turn reuses the same ses_ (session hit, depth=2).

Tool-call loop

Upstream answers with opencode tool calls (read, bash, …) that the client never asked for and cannot run. Forwarding them as-is kills plain chats (empty reply / error), so the proxy loops instead: every call is answered with "tool unavailable, answer directly" and the request is resent (same session, fresh msg_), until clean text arrives. usage is summed across rounds. If the model keeps calling tools past MAX_TOOL_ROUNDS, the client gets the accumulated text, or a short note if there is none — opencode tool_calls never leak to the client.

MAX_TOOL_ROUNDS=0 ./target/release/oc  # single attempt, no follow-ups

Forwarding client tools (FORWARD_CLIENT_TOOLS=1)

Off by default. When on, client-declared tools are appended upstream next to the template (verified: extras don't 403, verified on both chat/completions and responses), so the model can actually call them. A round that calls only client tools is returned to the client as-is (tool_calls + finish_reason: tool_calls, streamed as deltas) — the client executes them and sends back role=tool results, the loop continues upstream. Calls to template tools (or unknown names, or mixed rounds) are still stubbed as above, but with forwarding on the stub then names the client tools the model may use instead (with forwarding off the classic message is kept, so the model isn't tempted to hallucinate calls to tools it cannot see). In NixOS: services.oc-zen-proxy.forwardClientTools = true;.

Proxy failover + retries

Every outbound path (direct + each proxy) has its own cooldown: a 429 parks that route for RATE_LIMIT_COOLDOWN_SECS (default 120s, Retry-After wins when present — upstream FreeUsageLimitError carries no timing itself), a transport error or 5xx parks it for ROUTE_ERROR_COOLDOWN_SECS (default 30s). Each logical turn is sent up to MAX_UPSTREAM_ATTEMPTS times (default 3) with exponential backoff (RETRY_DELAY_MS, default 800ms), skipping parked routes — so a rate-limited direct is followed by proxy attempts within the same client request, and a route that failed minutes ago is not retried until its cooldown expires. When every route is parked, the turn does one live parallel recheck of all egress and uses whoever answers; only if nobody does, the client immediately gets the last real upstream 429 body verbatim, with no further waiting. Retried: transport errors and 429/500/502/503/504. Permanent errors (400/401/403/404/422, incl. geo-block 403) return immediately.

A counter of consecutive retryable failures at PROXY_FAILOVER_AFTER switches the preferred order to proxies-first for PROXY_COOLDOWN_SECS (default 180s): proxy successes keep the window open, the first direct success resets the counter and returns to direct. Switches are visible in logs (route rate-limited, cooling down / failing over to proxies / retrying upstream request ... / back to direct, passwords never logged). /v1/models always goes direct (it has a static fallback).

Egress health

Every outbound path (direct + each proxy) cools down on failures: a 429 parks that route (RATE_LIMIT_COOLDOWN_SECS, doubled per repeat up to RATE_LIMIT_MAX_SECS), a transport error or 5xx parks it briefly (ROUTE_ERROR_COOLDOWN_SECS). Limits are per model and IP, so probes run one inference request (template + tiny cap) per PROBE_MODELS model across both wire formats, and a route counts as healthy only if every probed model answers; /models alone is never probed, it is not rate-limited and reports healthy on dead IPs. Probes run on manual /v1/update and automatically as a live recheck when every route is parked; traffic successes unpark at once.

Manual refresh: curl host/v1/update deep-probes every egress in parallel, one IP MODEL STATUS line each, no credentials: 138.99.37.196 nemotron-3-ultra-free 200. Dead routes park themselves, recovered ones rejoin rotation at once.

Pool status without probing: curl host/v1/health lists alive egress (one host per line, parked ones omitted).

export PROXIES="socks5h://u1:p1@1.2.3.4:1080,socks5h://u2:p2@5.6.7.8:1080"
export PROXY_FAILOVER_AFTER=2
./target/release/oc
# or via file: PROXIES_FILE=/run/secrets/oc-proxies

Secrets outside git and the store

proxiesFile/apiKeyFile in the module are plain strings, not nix paths: the file is read at systemd runtime and never copied anywhere. Keep it e.g. in ~/secrets/socks5 — no need to commit it, no --impure needed:

services.oc-zen-proxy.proxiesFile = "/home/lily/secrets/socks5";

Format: one URL per line, # comments. Two requirements: the file must exist on the target machine and be readable by the service user — if home is closed off (chmod 700), allow traversal: chmod o+x /home/lily /home/lily/secrets && chmod o+r /home/lily/secrets/socks5 (or ACLs via setfacl). For production, sops-nix/agenix with files in /run/secrets/ is the right call.

In the NixOS module it's even easier: set proxiesFile and the rest happens by itself — on every nixos-rebuild switch the dir and file are created if missing (template inside) and chowned to the service user. Content of an existing file is never touched. Uncomment the lines, add your proxies, then systemctl restart oc-zen-proxy.

NixOS

nix build .#default

Wire up the module (flake input oc) — the package is substituted automatically, no overlay needed:

{
  inputs.oc.url = "path:/home/lily/Documents/oc"; # or git URL
  # ...
  imports = [ oc.nixosModules.default ];
  services.oc-zen-proxy = {
    enable = true;
    port = 1234;
    host = "127.0.0.1"; # localhost only, nginx faces outward; 0.0.0.0 listens everywhere
    apiKeyFile = "/run/secrets/oc-api-key";   # or apiKey = "..."
    proxiesFile = "/var/lib/oc/proxy";        # auto-created with a template if missing
    failoverAfter = 3;
  };
}

Public access via plain nginx: https://ai.example.com/v1

The module doesn't manage nginx — just drop a standard virtualHost next to it:

services.oc-zen-proxy = {
  enable = true;
  host = "127.0.0.1"; # localhost only, nginx faces outward
  port = 1234;
  proxiesFile = "/run/secrets/oc-proxies";
};

services.nginx = {
  enable = true;
  virtualHosts."ai.example.com" = {
    forceSSL = true;
    enableACME = true;
    locations."/" = {
      proxyPass = "http://127.0.0.1:1234/";
      proxyWebsockets = true;
      recommendedProxySettings = true;
      # SSE streaming + slow free-tier turns (tool rounds can take minutes).
      extraConfig = ''
        proxy_buffering off;
        proxy_read_timeout 300s;
        proxy_send_timeout 300s;
      '';
    };
  };
};

Check: curl https://ai.example.com/v1/models.

All options are in nix/module.nix.

Limitations

  • ~10k prompt tokens of overhead per request — same as opencode itself sends, the proxy adds nothing on top (and even saves the TUI title requests). But unlike opencode, LM Studio doesn't compact history — long dialogs hit the context sooner.
  • Model persona is opencode (coding assistant), not neutral.
  • Reasoning deltas (nemotron) are folded into reasoning_content in non-stream responses.
  • LM Studio client tools are not proxied (exact opencode set required). Upstream tool calls are auto-answered as unavailable (see "Tool-call loop"), so coding questions get a text answer instead of workspace actions.
  • Each tool round re-sends the growing history: multi-round turns cost extra prompt tokens.
  • Free models sometimes answer 503 Upstream overloaded — retry. The free roster rotates, some have geo-blocks.
  • Template is hardcoded: if opencode updates system/tools/UA, expect 403 until assets are re-captured.