Show HN: We built open OpenRouter that turns usage into a better model

220 points
1/21/1970
3 days ago
by SilenN

Comments


Areibman

Could you say more about how caching works? One major advantage of sticking with a single model is saving money on cached input tokens. I'd imagine if you swap between a bunch of models, you may improve performance but cost would would balloon out of control

3 days ago

SilenN

The trick is to rarely switch, or switch at task boundaries. Often the conclusion of routing is actually "this one model is actually at the pareto front for this task, just use it always".

3 days ago

cameronh90

But then it's better to just not have a gateway switch models at all.

Just have the harness able to choose which model its sub-agents use, then tell it how to split up tasks and which models to use when doing so.

3 days ago

SilenN

That is another way to do. Or we can automatically figure out which models the subagents should be using for you. And update them as new models come out and the work your subagents do changes. More than one way to skin a cat.

3 days ago

aerzen

This does make sense. I generally only switch between models in pi when creating a new session. And it is apparent from the promt if this just a "how to see open ports on linux" or "make a concrete plan for feature X"

3 days ago

try-working

Generally you should only have two models in the pool per domain. I wrote some of my learnings building a router here: https://try.works/first-principles-of-model-routing

3 days ago

purplecats

and caching is related to performance too ofc

3 days ago

mongrelion

Congrats on the launching of your product. I will be taking it for a spin to compare it with these other products that seem to be competing directly with what you have to offer:

- https://github.com/ENTERPILOT/GoModel - https://github.com/maximhq/bifrost - https://github.com/BerriAI/litellm

Would you care to share what makes experiential different?

3 days ago

jakswa

I think I have to look at setting up one of these AI gateways for work, since AWS bedrock is such a PITA to hook a harness up to over IAM roles. Also I still can't believe bedrock hasn't released any open models in months (so there's paranoia that I'll want to swap in another provider).

Really though I'm hoping anthropic fixes the oppressive claude verbosity. I saw someone refer to being "clauderboarded" and my brain cannot let go of this as Claude's tokens bombard me.

3 days ago

SilenN

We had this exact problem so we solved it for ourselves. Happy to help if you run into any issues.

I have a /hmmm command I use for the second part that works reasonably well: "Stop using jargon and speak coherently. State it more simply and concisely, like one human talking to another. Make it like google dev docs style. More dead prose. No aphorisms, no flourishes. Simple."

2 days ago

jakswa

Oh. The UI screenshot on github is... not actually in the github repo? It's platform/hosted only? There's my first awkward discovery, but makes sense in retrospect.

2 days ago

croemer

Probably shouldn't call it "Open router" in the title as that's a specific brand. Maybe you meant "we built something like OpenRouter".

Rereading I see you wrote "open OpenRouter", which looks a bit like a typo at first glance.

3 days ago

nejch

A part of me wishes the open source community would focus making research and industry-backed initiatives like the vLLM Semantic Router rock solid. Then I'd spend less time every month checking if this or that new model router has differentiating over vllm-sr :)

At least for open source inference, it seems like there's healthy competition centered around vllm/sglang, but 2026 seems to be for model routers what 2025 was for agent harnesses.

3 days ago

akshay_akula

Open source and no markup is the right default for a gateway. The caching question above is the one I would want answered before swapping models though.

3 days ago

SilenN

Ans: we rarely switch, often times it's just a "switch to using this model for your agent"

3 days ago

akshay_akula

You guys should look into ngrok ai gateway. We have some small models running on local hardware that we tried to use but it was too painful. Just wanted to not waste all the compute hit one endpoint and be done but as a team.

3 days ago

ceroxylon

>The gateway adds under 1 ms for BYOK requests

Amazing! Really brilliant idea, thank you for sharing this project. There is so much ground to cover in the LLM gateway / routing / reporting world, and this is a great start. The Tinker implementation is my favorite part, fine tuning is much better than a sea of context files.

3 days ago

kfallah15

Thanks! We are going to add continual RL via Tinker soon too

3 days ago

cheema33

I have not tried it yet. Is it similar to LiteLLM? If so, what sets it apart?

3 days ago

kfallah15

Router and model optimization from traffic is the main differentiator

3 days ago

SilenN

Also a hosted marketplace, not just BYOK

3 days ago

foremerge

We built our own model router at GPTree but this looks interesting. Caching is definitely one of the hardest parts to get right, especially in our scenario where users can branch from any part of the conversation and keep the dynamic context from the parent (unlike most other platforms that lock in the context once you branch).

3 days ago

23david

Super interesting and congrats on the release. Curious if you initially had this in Python and then rewrote in Rust?

3 days ago

SilenN

Yep! If you look at the commit history that's exactly what happened.

3 days ago

0xbadcafebee

You started it a week ago? I look forward to checking back in 3 weeks when you've exited for $1B

3 days ago

tyre

Looks like first PR is June 24th: https://github.com/experientiallabs/experiential/pull/1

So, two months. Still impressive!

3 days ago

SilenN

Thanks for the positivity tyre! If you look at our git history, we pivoted and only started building the gateway recently. Before that we were building research infrastructure that now powers the intelligence features we provide.

3 days ago

rdslw

impressive only if using pre-gpt era assumptions about saas/products/software.

unfortunately a small team can reproduce it in two months, which greatly lowers value of it.

we, as a collective, have to change our value-judging logic and tune it to post AI world.

3 days ago

SilenN

See you soon

3 days ago

d2p

> and use your traffic to (opt in) train you a model.

Is there more info on this? I'm curious exactly what it is. Is it fine-tuning/LoRA on some base model? Don't cloud providers encrypt reasoning now - does that prevent this?

3 days ago

forgetme2020

what's the business model here. How does experiential labs make money

3 days ago

kakugawa

They make money on enterprise plans: https://www.experientiallabs.ai/pricing#enterprise

Look at the Intelligence features in the Enterprise plan:

* Per-prompt model optimization

* Caching

* A model you own, trained on your traffic

3 days ago

kfallah15

yep, it will be through enterprise licenses and our own hosted platform built on the repo

3 days ago

rdslw

the business model, I suspect, is classic rug-pull in some time after building user base.

proof: boldly claiming being open source in literally first sentence, while cowardly hiding on-by-default telemetry (WTF??) in truly last paragraph of readme.

sorry to sound harsh, but this is typical old era playbook here.

in the era of AI, fortunately, such products has much lower value. people and VCs didn’t yet tune to it.

3 days ago

SilenN

Telemetry is off by default. PostHog is for usage analytics on the open source repo. Audit it if you're skeptical.

We make money off enterprise licenses and hosting models.

I have strong reason to suspect you either can't read or are a bad actor.

3 days ago

aHumbleUser

Your GitHub readme says telemetry is on by default

3 days ago

SilenN

"Anonymous aggregate PostHog product telemetry is enabled by default. It never includes prompts, traces, actions, observations, paths, model names, credentials, or raw customer content."

2 days ago

swthbht

Very cool. Does your gateway decide effort levels as well? Or just models?

3 days ago

SilenN

Yep! One interesting example is often Opus 5 on low reasoning ~= Opus 5 on high reasoning.

3 days ago

sangwook

What online signal recalibrates simulated rankings against actual task success? Also do you have a plan to support semantic caching at the router level?

3 days ago

kfallah15

For the online signal, we use a LLM judge with a rubric calibrated offline by the user via TUI. UX of the calibration is a major focus area. Semantic caching is interesting, open to supporting it but not currently planned.

3 days ago

bicepjai

Great tool. Thanks. I am working on something similar.

2 days ago

carloslfu

why text world models? Is this backed by any evidence or did you find it to be good in practice?

2 days ago

ashermania

Finally an open source tool doing this!

3 days ago

gpiechnik2

great design! i love it

3 days ago

SilenN

Thank you!

3 days ago