d@n tech


Caffinated Tech Insights


the gooseneck kettle grew a barista. last time i wrote about gophergate, it was a clean little proxy: requests go in, model alias maps to a provider, response streams back, spend gets logged. boring on purpose. that was the whole point. but a few months of running every ai call in the house through one box changes a thing, and gophergate stopped being a reverse proxy and turned into the front of house for a small café.

the model pile

the provider list grew from five to seven, and the models behind them exploded. i’m routing across:

  • openai, now up to gpt-5 and gpt-5.4, plus the o-series reasoning models and dalle image gen.
  • google gemini, including the gemini 3 flash and pro previews, with imagen 3 for images.
  • deepseek, chat, reasoner, and the v4 flash and pro models.
  • moonshot, kimi k2.5 and the corrected k2.6.
  • xai grok, grok-3, grok-4, grok-4.3.
  • xiaomi mimo, v2.5, the one that surprised me most.
  • ollama, for the local models on the homelab.

that’s a lot of beans to keep straight, and the naming churn is real. kimi was kimi-2.6 in one place and kimi-k2.6 everywhere else, so i fixed that. gemini-3.5-flash-lite turned out to be a name that didn’t quite hold up, so it came out of the targets. and reasoning_effort=none is an openai-only thing, so it’s scoped there now instead of every provider rejecting it. the details are the whole game.

routing stopped being a lookup

the first version of gophergate did a lookup: alias to provider, done. now it has actual routing strategies and it’s a lot more interesting:

  • hierarchical routing lets groups point at other groups and cascade until a concrete model is reached, with cycle detection and a depth cap.
  • heuristic routing is the free, instant one: keyword and condition checks against tags, token limits, multimodal inputs, reasoning, and tool calls.
  • classifier routing uses a cheap llm to rate the task on a 1-10 complexity scale, then picks the model bucket. it’s the “how hard is this really?” strategy.
  • two-level dispatch runs a dispatcher group that reads the complexity score and hands off to tier groups, each with its own internal strategy.
  • and it’s provider-aware, so classifier selector models route to the right provider automatically instead of bouncing off a wall.

the gooseneck kettle now has a barista who smells the beans before deciding which to grind.

the dashboard

this is the part that stopped feeling like tinkering and started feeling like a product. there’s a management dashboard now, with usage analytics, per-request cost tracking, and per-day activity. i added a days active card and the total lifetime span, and redid the stat-card grid so it reads clean on a laptop without collapsing into a mess on a phone. there’s a mobile off-canvas sidebar now, so i can check the numbers from the couch.

roles matter here. there’s an admin role with full access and a viewer role that’s read-only on usage and cost. plus client api keys, so external integrations can hit the proxy without me handing out the master key. i can see exactly what every request cost, in real time, down to the token.

boring reliability, dressed up

beneath all of it, the same philosophy: make it boring so i’m not paged at 3am. that turned into a specific engineering pass:

  • connection pooling with a shared http transport, max idle connections up to 200, tcp keep-alives, and http/2 multiplexing across providers.
  • in-memory token caching with a sync.map, a 10 second ttl and a 2 second negative cache, so auth doesn’t hammer sqlite on every request.
  • thread safety with rwmutex locks across the provider maps, model registry, and router reloads.
  • circuit breaking so a dead provider gets isolated and recovers on its own after a timeout.
  • async logging to sqlite from background workers, so the request path never blocks on a write.
  • security: hmac-sha256 signed session tokens for the dashboard, encrypted api keys in the database, and token redaction in logs so a secret can’t leak out through a ?key= in a url.

and since the last post, deepseek got openai’s responses api. the /v1/responses endpoint now works for both openai and deepseek, which is a real chunk of new capability for a workday.

where it’s going

it’s close to presentable. the codebase is tidy, the dashboard is genuinely useful, and the routing is the kind of thing i’d hand to a friend. the open source push is still on the list, and increasingly it’s a “soon” instead of a “someday.”

the ai model landscape is a moving target, and this month’s hot model is next month’s footnote. but the layer underneath, the proxy that routes, tracks, and fails over, that’s stable. you can swap beans all day and the shop still runs.

i need to go re-check the classifier threshold. the coffee’s fine, but i think the barista is over-tipping.

-dustin