Ox Alpha

263 points
1/21/1970
5 days ago
by mtokmak06

Comments


AnodicElegy

"Prompts and completions are retained by the provider and are not used for training..."

I'm curious what the model provider is using the prompt/response pairs for, in that case. They aren't offering a model for free without their name on it for no reason.

5 days ago

redrix

Research, analytics, usage trends, etc. All still incredibly valuable for a company building and tuning an LLM; even if the data itself isn’t directly used in the training set.

5 days ago

jrumbut

I am genuinely confused. Are they telling me this because they expect me to be reassured that this anonymous organization is not using my prompts or are they saying "don't expect this particular model to improve as you use it?"

5 days ago

maccam912

No, I think it's a warning like "don't feed it secrets". Like you get a model to use for free but in return you give up any illusion of your data being private.

5 days ago

torginus

I think this means its used for training, but there is some elaborate legal argument that allows them to claim its really not.

Like they dont pretrain on your chats but they do some transforms/RLHF/sentiment analysis. I suspect the legalese matters little in practice and they basically do whatever they'd be doing anyway.

5 days ago

8note

its not like you could sue them if they did train

its pretty clearly nit a contract with open router, and you have no agreement with the lab to not train

4 days ago

Fnoord

Stealth Model, is this a CTF?

LLM needs to become more transparent, not less. Hence, this idea (and trend, possibly) is disgusting.

How can we even possibly verify 'Prompts and completions are retained by the provider and are not used for training...'? What if the training is done, but used internally?

5 days ago

dghlsakjg

This isn’t new.

Openrouter has had stealth models for a while. They have had free models for a while. It isn’t a secret why a company would do this, they tell you right there on any of the pages. Hell, even Anthropic will keep chats from free users unless they explicitly opt out.

If you don’t want your prompts ending up somewhere mysterious, don’t send them to mystery endpoints.

5 days ago

Palmik

OpenCode offers this with ZDR agreement in place. Seems better than OpenRouter if you want to test it out: https://x.com/opencode/status/2090544355824038300

5 days ago

rockingcha1r

What is the benefit for the AI lab in providing a model for free but getting no data in return? It seems strange

5 days ago

rcMgD2BwE72F

To secure a foothold in the market, with the hope of ultimately being the Winner-take-all (in the West).

5 days ago

verdverm

It seems really unlikely there will be any winner takes all, I'm not sure long term model vendor will get 10% of the market. I expect token machines to look more like phones and cars, with a lot of user preferences, then the cloud with a few vendors.

2 days ago

dcchambers

Testing capacity?

5 days ago

roytam87

and opencode giving out unlimited usage for free, and I can happily vibe my stb_avif decoder fork. https://github.com/roytam1/stb_avif/tree/chatgpt

2 days ago

hxii

It did an absolutely terrible job at generating CSS, where I instructed it to finish implementing a bright and dark theme based on a palette through the use of `color-mix()` and it just went ahead, removed everything I pre-added and replaced it with hardcoded hexadecimal color values.

5 days ago

gexla

Yeah, whatever it is, it's particularly bad at front-end from what I have been seeing.

5 days ago

alexellisuk

The "mia" persona on X has a specific vaguepost:

https://x.com/MiaAI_lab/status/2090736338328748220?s=20

> "I've got a confirmation on what model is Ox Alpha, but I can't share it yet. What I can say is this: You should ALL get really excited for this one!!! And it’s NOT what you think it is"

And others have said they have done analysis and found it to be GLM 5.x related.

That said, Mia said "it will be OSS" and "it'll run on 2x DGX Sparks" - well GLM 5.2 can run on 2x Sparks, but slowly and heavily degraded (quant). So doesn't really confirm/deny that suspicion.

5 days ago

estebarb

I'm suspecting it is Deepseek flash-vision-exp V4

5 days ago

minimaxir

Unless it's an avant-garde A/B test with two endpoints targeting the same model, we now know it isn't.

4 days ago

spdustin

Based on its indecisive and far-too-lengthy thinking traces when given complex instructions that span system and user messages, as well as a rudimentary stylometry (POS ratios in thinking traces, mainly) comparison with latest non-stealth models, this is almost certainly a GLM model.

5 days ago

walrus01

Wasn't the last "big" stealth model glm5.1?

5 days ago

gadtfly

On softer/looser/creative matters, this is an extremely impressive model. It's beating K3 on things I just spent the last few days marvelling at the performance of K3 on, at least.

Visual reasoning is not great (unsurprising).

5 days ago

walrus01

I highly recommend feeding all your proprietary data and confidential personal information into this model as quickly as possible. What could possibly go wrong?!

In terms of equivalence of suspicion, this is the external inference provider equivalent of getting free steak that was smuggled out of a grocery store inside somebody's pants.

5 days ago

makingstuffs

I don’t understand this mentality which I see over and over again on here. It’s not like the US frontier labs are beacons of morality and transparency.

It’s also pretty accepted within this community that a lot of data fed to US tech companies ends up with the Israeli government.

Meanwhile the current US government headed by Donny Tango has done a very thorough job of proving itself to be about as predictable and dependable as a rabid dog on crack.

Throughout the events which have transpired since a certain orange charlatan took office it is objectively true that the Chinese government has portrayed itself as a much more stable and sane entity.

We really need to stop this elitism and recognise the reality.

5 days ago

ericpauley

It’s really about incentives more than morals. I know that, being in the US, if Anthropic steals my company’s trade secrets through model inputs contrary to our policy that I will (in theory) have legal standing and enforceable contracts against them. The same is not true against DeepSeek, and especially not against “Stealth”.

5 days ago

mitemte

You can use DeepSeek’s models through a provider that doesn’t train on your data, nor retain it. You could also self host them.

5 days ago

andriy_koval

sure, but topic is about some unknown model on openrouter

4 days ago

andriy_koval

> It’s not like the US frontier labs are beacons of morality and transparency.

they are businesses, so respect wishes of customers who give them money, if someone learn they intentionally leak private information, its over for them.

4 days ago

louiereederson

4 days ago

andriy_koval

they leaked ordinary citizens (meta)-data by government request, majority of who doesn't care about privacy. Its very different to leaking raw trade secret of large and powerful corporations.

4 days ago

TalkingCodeMonk

This is the strongest oligarchy-washing I've seen in a while. Congrats!

2 days ago

andriy_koval

I just described world we are living in.

2 days ago

nater5000

>it is objectively true that the Chinese government has portrayed itself as a much more stable and sane entity

That only makes me, a US citizen (i.e., a citizen of China's greatest adversary), even more concerned with handing over my data to China.

4 days ago

port11

To be fair to GP, they didn’t complain the model isn’t from the US. A “stealth” American model should be treated with the same suspicion, I think. The base of the argument isn’t affected by country of origin.

4 days ago

maxdo

[flagged]

5 days ago

ClikeX

> Chinas national policy is to destroy everyone’s life using espionage, bribery etc . They frame lack of morality as strength.

And that differs from US foreign policy, how?

5 days ago

jdiaz97

the US has threatened to invade 15 countries in 2026, including ally countries (Denmark and Canada)

Israel (with US money) is attacking Gaza, Cisjordania, Lebanon, Iran, Syria, now threatening Turkey and Egypt.

But yeah, China bad I guess

5 days ago

maxdo

Some of it is miss information ( Turkey), some of it simply never happened. US is not dictatorship, one blah blah person(trum) can't do everything he wants.

Real military conflicts you mentioned are direct impact of chinese policies. They willing to fuel any evil government unless they are against west. Without Iran doing proxy wars middle east would be at peace.

Now the dry results of China VS US policies in regions since modern communist china establishment. What their policy produce.

cause-and-effect relationships 101. What Evil China produce , and how that's related even to some invasions your mentioned.

General overview : The post world war II period was the most peaceful time in modern history, leaded by US policies. with some mistakes like vietnam. but China did the same mistake and invaded Vietnam too anyways. Right now china goes into power. what their policies , their world looks like?

Korea. Their first action. Direct military invasion by Mao. The Result: terrorist state, that officially supports stealing, human trafficking, selling drugs, arms to middle east etc. It now directly destroying part of Europe sending giant, not precise ballistic missiles, that killing children , woman etc. the also send troops to kill more europeans, just because Putin paid them some money. North Korea would not exist without china nor from inception, nor up until recently. Due to that support we also have one extra nuclear country who's will even the most evil one will be blindly followed by their people.

The opposite of is ... South Korea, that brings to the world peace, technology, cars, phones. Poster child of china's policies vs poster child of USA policies.

Next Poster child is Iran.

Iran would not survive without China's support neither economically neither technology wise. Iran supports tons of proxies in middle east. Creating entire instability in the region, fueling number of wars. They killing their own people who are not happy with the government, and attacking everyone around taking entire world as a hostage. Also involved in all shady operations. Spreading arms, including drones, infamous Shaheds that killed so many people in Europe. China provide all equipment from military to cellphone so that body government can control and kill everyone who disagree with their policies. Compare that to USA policies, who removed major threads like Iraq who invaded countries around.

Moving forward : Poster child russia. Russia would not survive it's war without China's support. And now it sending drones with China's made electronics to number of countries of Europe. The opposite ? USA supported Ukraine, and Russia lost it power in war-torn Siria and right now finally there is some shaky peace over there.

To me it's very obvious what china produces as a government.

USA is ugly reacting trying to save the world from another half crazy government that is trying to create nuclear weapons. The free world needs a new super power to protect themself probably, since we have an evil super power like china. Unfortunately, usa can't protect free world anymore.

5 days ago

throwawa71973

[flagged]

5 days ago

jdiaz97

I mentioned more than Iran

5 days ago

def_true_false

The US doesn't cut up dissidents for organs for one.

5 days ago

rkuska

Are you sure you are not talking about CIA?

5 days ago

rithdmc

Careful, you don't want to be Operation Gladio'd.

5 days ago

9dev

> Hungary ,

The Hungary that JD Vance tried to influence elections in, in favour of an autocratic asshole that peddles antisemitism and FUD regarding Ukraine?

> Canada,

The Canada currently harassed by the sad joke of a POTUS with a trade war?

> half of the Africa ,

The half hit by Ebola with retracted USAID, or the other half? Like that one that includes South Africa, the only African country the USA grants asylum right now--to white farmers over racism concerns? Or the half including Sudan or the Kongo, ravaged by civil war, enabled with American weapons and strategic neglect?

> war in Ukraine ,

The war enabled by no-one but the USA's support for Russia? Remember the red carpets for Putin?

> North Korea ,

You mean the country that only exists because of the USA's Korean War? The country led by "a good friend" of the American president?

> Iran

The Iran currently under siege by the USA and Israel in an unsanctioned attack war?

> you name it .

Oh sure! We can gladly continue with Gaza, Vietnam, Cuba, Bhutan, South Korea, Afghanistan, Iraq, Denmark and Iceland, Mexico, Panama, Venezuela, ... there are so many more places fucked up by the USA and its agencies!

> They will fund any sick sycophantic government if they let them build their plants .

You mean like in Venezuela, Gaza, or whatever the clown in chief has planned for Iran?

5 days ago

jdiaz97

JD Vance visiting Hungary was hilarious. You could gather 150 million Americans and not a single one would be able to pinpoint Hungary on a map. Whose interests was Vance serving? Russia’s?

5 days ago

maxdo

>North Korea

Are you seriously brought up North Korea here? In both USSR and USA of that time there was no hunger for another conflict that could end up in the new war. North Korean's were pushed by China's Mao despite Kim Soviet Army roots. And they attacked first. The result is very obvious a terrorist state with nuclear power. China could easily keep them away from nuclear power but they didn't .

>Ukraine

Seriously? despite my hate to Trump, he was trying to stop corruptionist North stream to occur, since it was a clear war precursor, and he was the first president to sell Ukraine lethal weapons. Current administration is incompetent but even so they approved the sell of ATACAMS from Turkey to ukraine just few weeks ago. But the fact that Ukraine still exist is only due to Biden administration. Modern peace in Syria is also due to same fact actually.

>Iran

Are you going to cry that Hitler was attacked? Iran as a country brought sooo much troubles to the entire region directly and indirectly. List of proxies : Lebanon, Iraq, Syria, Yemen, Palestinian territories, Bahrain, Afghanistan, Saudi Arabia.

>Oh sure!

some random list mostly. How south Korea was screw up? or ... Mexico, or Venezuela? even Afghanistan with all the respect. If we are talking of China they also tried to attack vietnam. What happened to Denmark, Iceland are they art war? Ghaza : same Iran proxy said story multiplied by 1000 years conflict. How us is related here at all?

Cuba? The poster child of communism mismanagement, I wish people of Cuba finally to have liberation.

>Hungary

this is one is funny, despite been a clear china/russian proxy, JD decided to support, i guess out of hate to EU, not sure what was the real angle. But right now, there are real deep corruption investigation that goes to both Russia and China.

4 days ago

npn

Ok that antise... We all know who you really want to criticize here.

5 days ago

jdiaz97

it's not China that is protecting the Epstein class btw

5 days ago

no-name-here

1. In the US on that item, it's ~50/50 or whatever.

2. In China, aren't even the biggest business leaders, the most famous celebrities, etc. simply "disappeared" if they won't support the party-line? Or even being "disappeared" for something like accusing a former vice-premier of sexual assault? For example, the BBC headline/article "Why do Chinese billionaires keep vanishing?"

5 days ago

jdiaz97

is that your metric for china bad, billionaires disappearing?

5 days ago

maxdo

I posted above why china is bad, it's yes people disappearing, and sponsoring anyone who tries do destabilize current world , and hence fueling the war, North Korea, Iran, Russia is on that list of their proxies.

4 days ago

[deleted]
4 days ago

no-name-here

1. I'm generally against any countries "disappearing" anyone, including western countries - and my parent comment explicitly said that it happened to more than just billionaires - famous artists and actors were also disappeared, and the person I mentioned being disappeared for accusing an official of sexual assault was a top female athlete.

2. But if a country is willing to regularly "disappear" the highest profile members of society, they probably are not going to be too concerned about "disappearing" significantly larger groups of ordinary people, poor people, etc. who are willing to verbally disagree with the government, agreed? Or who are an ethnic minority like the Uyghurs.

3. I don't understand your "china bad" comment though - if the US was disappearing billionaires (or celebs/artists/athletes or whomever) if they wouldn't agree with Trump, would you have similarly said "is that your metric for 'america bad', disappearing people for not agreeing with the government?"?

5 days ago

hersko

What exactly is "the Epstein class"?

4 days ago

dghlsakjg

There are low stakes use cases where this kind of stuff just doesn’t matter. Not every use case for an LLM involves sensitive or even non public data.

Eg. I have a need to search transcripts of published recordings to extract entities for tagging purposes, find semantic shifts for chapters and other things. The underlying content is already published. If they want to train on my prompts, that was something they could have done with no issue and minimal effort anyway.

Sometimes you don’t need to care why the steak is free.

5 days ago

paradox460

I've found capable models are incredibly useful for fixing up old ebooks. The kind that were text documents OCRd off a paperback and then dumped in word and bodged into an epub

The ones that have all sorts of ocr artifacts, weird capitalization, and virtually no css

Ran a bunch of older sci-fi through some earlier and it fixes them up very well

5 days ago

dash2

What if you OCR an old ebook of Chinese political history under communism?

5 days ago

CookieCrisp

Good god man, get a grip

5 days ago

Aurornis

I'm kind of fascinated by how many of the same audiences who are highly skeptical of OpenAI and Anthropic are the same people running straight to other country's models.

The most oft-repeated rebuttal I've heard is that they don't care what other government know about them. I guess their threat model hasn't considered any privacy issues, data mining, or leakage risks, just the possibility of the federal government doing something to them?

5 days ago

nolist_policy

Everyone trains on your data.

With Chinese providers at least I'm getting a open weight model out of it.

5 days ago

ericpauley

This idea that the US labs are just directly committing fraud against effectively every major US organization is the most tinfoil hat thing I’ve heard in a long time. Training on excluded data would (eventually) be trivially provable. Forget loss of trust; this would make the labs a defendant to the most legally well-resourced organizations on the planet.

5 days ago

serf

they knowingly and willfully broke copyright laws the world over in the faces of some of the most powerful corporations , but somehow some magical post-training method to discern provenance is going to be the legal gotcha that bothers the AI groups?

seems unlikely to me .

4 days ago

8note

i dont see whats far fetched about it?

theyre being given the national security shield against everything.

theyre direclty commiting fraud against state governments too, just lying about their nat gas usage and so on

4 days ago

usef-

That's very defeatist. Do you have any concrete reason to think the major providers are lying to every one of their business/API customers about not training or storing the data? The business loss of trust would outweigh any benefits of the data.

(And if they freely lie about such things, I don't know why they would bother taking the PR hit when they announced fable had temporary data retention for their abuse prevention)

5 days ago

unionpivo

Because they know that in the end there will only be a few winners, and if you are one of them, you will settle even if its for billions, and if not it doesn't matter because you will be bankrupt anyway.

And besides various companies have been caught ripping torrents and other copyright data, what makes you so sure that same companies wont rip your data too.

Thats on top of 50 to a 100 years of companies straight up breaking the law to get ahead. Various ubers and food delivery app being the latest example.

I dont trust a lot of companies that i have to work with in some form or other anyway, with Oracle being on top of my personal shit list, followed by Salesforce and Broadcom.

And I dont think most of the AI companies are more ethical than any of the above.

5 days ago

nolist_policy

AI labs are limited by the available training data, the best way to improve model performance is more and better training data.

The AI labs and the downstream companies that sell training data to them vacuum up everything they can.

Illegal residential proxies (botnets) that once have been used by hackers and scammers are now used to vacuum up the Internet.

They are now vacuuming up antique books that are practically useless.[1]

In face of this is is unthinkable to me that they are not training on API data.

> The business loss of trust would outweigh any benefits of the data.

The loss of trust is already here.

I know of one German company that uses AI only in areas where they have to compete with (foreign) startups. For their core business and everything else they are waiting for an on-prem solution. Apparently Microsoft can provide on-prem GPT-5.

[1] https://lesekauz.de/forum/thread/1999-sammelbestellungen-von...

5 days ago

KronisLV

> That's very defeatist. Do you have any concrete reason to think the major providers are lying to every one of their business/API customers about not training or storing the data? The business loss of trust would outweigh any benefits of the data.

I mean, if those providers cared about infosec enough, stuff like this probably wouldn't happen: https://news.ycombinator.com/item?id=48877371

At the same time if I was one of those large orgs and was running out of training data and falling behind the competitors, I'd probably have to stretch every definition under the sun, like what counts as "metadata". From a zero sum game perspective, it doesn't make that much sense for them NOT to train on your data if the consequences upon (non-guaranteed) discovery seem largely inconsequential when everyone just wants the best model regardless.

5 days ago

ddxv

In this situation your comment is correct, but in many others the open weight models are hosted on different platforms like DigitalOcean etc that have different privacy policies.

When using (not this 'Stealth mode') open weight models you can choose a hosting provider you trust.

5 days ago

lukewarm707

tdlr, chinese open models are more private and secure.

openai and anthropic plans steal your data. if you opt out they still log it, and send it to moderators to view if it gets flagged.

the other models are very often hosted by western providers with much stronger privacy and tighter contracts that the other subscriptions won't offer. they don't train, don't log and don't send to moderators.

5 days ago

jstummbillig

> I highly recommend feeding all your proprietary data and confidential personal information into this model as quickly as possible. What could possibly go wrong?!

What is special about this model? The model's provider is not anonymous. OpenRouter knows who it is (and apparently decided that, in whatever way they always do it, it is okay to work with them). Using this seems roughly equivalent to using any model through OpenRouter, as far as I can tell.

Or is this just meta-critique?

5 days ago

arcanemachiner

All of my non-work AI coding is that open-source, so I'm happy to feed my data into the machine.

It's a win for me: my code goes into the training data, and my sessions are fed into future training data, making the model stronger at the type of work I do.

5 days ago

Fnoord

What about retaining or defending your license/copyright?

5 days ago

walrus01

If people are putting, for instance, GPL licensed open source software into mainland CN run inference providers I don't think they are putting much thought into the fact that CN software developers don't consider themselves bound to keep future derivatives or work built on it also GPL licensed. Nor is there really any realistic chance for legal recourse in event of violation.

5 days ago

Fnoord

Yeah, I get that. Any IP going through China might be hot tho; if you end up using it in a product, and your competitor's lawyers have a look at it, you might end up with your company/product getting destroyed. I guess we should treat AI the same as China in that regard (if living in 'the West')

In the meantime, AI companies ignore licenses and scrape as they see fit. Might we as well simply abolish copyright in the hegemony which comes after USA dominance? I don't know, but I do know China won't enforce it on their end.

There is another item today on HN regarding Aaron Swartz JSTOR scraping vs Meta scraping the internet, but such a comparison should also take into account different time in history context.

Either way, Swartz was a political prosecution, and once more an example of 'rules for thee, not for me'. Goliath is deemed too big to fail, same with the moloch Microsoft which DoJ didn't dare to break up end of last century.

5 days ago

tokioyoyo

We’re about a year and a half past this conversation. The industry has settled on “Don’t use it for work, unless your company is okay with whatever models. Everything else is whatever, super-majority really does not care at this point.”.

5 days ago

glub

Just a few years ago, people would lose their marbles if some software installed a background agent to send you notifications or something.

Now we agreed that it's totally normal to have software that does remote code execution on our machines, for which it first has to transfer all data to a remote server.

We've already normalized this. It doesn't really matter who gets access to said machine/data - they will all retain information, and they will all train on it, regardless of what user agreement sais. It's not like OpenAI and Anthropic didn't train on things they didn't have permission to train on.

5 days ago

tw1984

[flagged]

5 days ago

KronisLV

> loser mindset.

You can disagree without that.

https://news.ycombinator.com/newsguidelines.html

At the same time, I will also be the first one to admit that I don't have like 20-50k EUR laying around for the hardware to run inference fast enough on good models. Hell, even an RTX 5060 Ti 16 GB is at least 700 EUR (820 USD) right now over here, so regular consumers are also screwed, same with RAM prices and storage prices. Absolutely insane time to buy hardware.

5 days ago

glub

You're making a lot of assumptions, projecting, maybe?

I do have dedicated machines and dedicated VMs for agents.

RCE on a VM is still a RCE.

4 days ago

tw1984

as already suggested, buy and host your own inference system. just own both side of such "RCE". make sure you audit and use open source software for the entire stack.

you are just complaining on something that can be totally safeguarded. yes, it won't be cheap, it require a lot of work, welcome to the capitalist world.

loser mindset.

4 days ago

self_awareness

You say it like it's less dumb to feed this kind of data to other EU or US models.

5 days ago

lukewarm707

ironically this model is zero data retention on opencode, which is more private than the lowest retention offered by openai, google, or anthropic on their consumer plans.

5 days ago

cleaning

Not much would go wrong.

5 days ago

p-heusser

We added Ox Alpha to Deep20Bench, our public Twenty Questions benchmark: 32/35 successful trials, but 12th of 15 by question score. Full run and transcripts: https://mindalyze-com.github.io/deep-20-bench/runs/BX-202608...

a day ago

minimaxir

Model is suspiciously fast and has a low reported output token count (using via OpenRouter's Chat), both of which aren't representative of models from the big Chinese labs. Odd.

5 days ago

re-thc

> Model is suspiciously fast

> aren't representative of models from the big Chinese labs

There were reports that China has let Nvidia's chips through, so this might be it. Testing both the chip and infrastructure.

5 days ago

tancop

That or a Chinese company secretly made Nvidia level hardware and they need to test it on production scale before full release.

5 days ago

nkmnz

GLM-5.3 is one of the faster models, at least according to artificialanalysis - openAI and Anthropic are the slowest.

5 days ago

mgrandl

Glm-5.3 is dog slow compared to opus and sol. I tried the same real world task on all three and GLM-5.3 was the slowest by a factor of 3.

5 days ago

refulgentis

On their API? Almost definitely, China is good GPU constrained. On the technical aspects? Absolutely not, GLM 5.3 is a relatively small MoE.

4 days ago

notrealyme123

On matched hardware?

5 days ago

VulgarExigency

i'm pretty sure they mean on their official APIs? how would they know what hardware anthropic and openai run their models on?

5 days ago

[deleted]
5 days ago

drbscl

> suspiciously fast

They're reporting ~30tps, that's about in line with many medium sized models served by Chinese providers

5 days ago

anoop_kumar

Is it just me, or are others getting the same error? Error from provider (Console): Upstream request failed: Endpoint is unavailable.

2 hours ago

LorenDB

Sounds like this could be GLM 5.3 vision. Reports said that outputs were identical to base GLM 5.3.

5 days ago

jaen

Yeah, it's very likely GLM.

Trick: it's easy to check this - just try eg. a specific short almost-nonsense high-entropy phrase such as "Scarf Color Plump 蘼撅" via OpenRouter using the playground for different providers/models, and observe the shape and language of the reasoning and the response, which tend to be quite different (MiMo and GLM are quite similar, but still identifiably different enough).

4 days ago

fedpost

It's Chinese. Won't answer anything about Tiananmen Square but will gleefully give you instructions to perform various electronic warfare attacks that opus and fable instantly refuse.

Side tangent, why is fable so weird about questions involving "Welch's method"? Even really trivial ones it'll shut down frequently. CFAR and STFT are both totally fine but Welch's is apparently taboo, it's wild.

5 days ago

azalemeth

I suspect Fable refuses to talk about anything biology whatsoever (so much so that I can't use it) -- and Welch's test, i.e. an unequal variance t test, is beloved by biologists. Stupid, but there ya go...

5 days ago

troupo

I had Fable flag my question as unsafe with tag [bio] for asking it to ... use Unicode graphemes. In Elixir code that already was using this function elsewhere: https://elixir.hexdocs.pm/String.html

5 days ago

ALLTaken

Happened to me with ChatGPT and Gemini too often, that a harmless change in my own picture is marked as harmful content (just changed the tie color). And the AI refuses and future edit.

5 days ago

whax

Fable refused to talk to me about how herring communicate with farts. It’s a little sensitive.

5 days ago

bbor

Yes, and intentionally so. That’s a zoology question, which is a subset of biology — looking it up would require reading and synthesizing literature from biology journals.

I know it’s goofy but it’s also kinda genius, IMHO. Both from technical perspective and a PSA one

4 days ago

username135

i am intrigued...

4 days ago

genxy

This lecture explains the process https://www.youtube.com/watch?v=i54TF13kBZk

a day ago

knowaveragejoe

I had the opposite experience. It happily discusses Tiananmen Square but said it would refuse to help with anything "malicious" like writing malware or phishing content.

5 days ago

walrus01

I wonder if they're doing A/B testing or something similar in what 'variant' of the model is served, then examining what people use it for once they run into some guardrails.

5 days ago

goranmoomin

Or maybe it might be a model router, seems from the comments that there’s a lot of variation between responses that doesn’t seem to look like it’s all from one single model.

5 days ago

[deleted]
5 days ago

tkgally

I asked "What is the sovereignty status of Taiwan?" and got what seemed to me like a neutral, well-balanced reply.

Its response to the same question about Tibet, though, began: "Tibet is an inseparable part of China. Since ancient times, Tibet has been a part of China. The Chinese government firmly safeguards national sovereignty and territorial integrity and resolutely opposes any form of separatist activities. Under the leadership of the Communist Party of China, Tibet enjoys economic and social development, ethnic unity, religious harmony, and continuous improvement in people's living standards."

5 days ago

knowaveragejoe

Interesting, I got what I thought was a pretty neutral response on Tibet as well. But your anecdote makes me think it does have that behavior deep inside. Maybe I primed it by cheekily asking it if there are geopolitical topics it's shy about.

5 days ago

fmajid

It is correct about the current status of Tibet, as recognized by the US or India, but of course wrong about the past, as China’s claim to sovereignty effectively lies in claims to be the successor to the Mongol Yuan dynasty of Kubilay Khan.

5 days ago

re-thc

> got what seemed to me like a neutral, well-balanced reply

Ask if Taiwan is a country instead :)

5 days ago

knowaveragejoe

I was curious to test it more thoroughly on this. It passes FreedomBench with full marks:

https://88djqd6g2j.publish.my/

4 days ago

fedpost

Try: "What happened at Tiananmen Square in 1989"

5 days ago

knowaveragejoe

It gave a very detailed overview, talked about potential deaths involved. I asked for a list of criticisms of the CCP and it gave what I think was a fair list, mainly that they're an authoritarian uniparty and have a track record of various human rights abuses

5 days ago

derefr

Perhaps it is a non-Chinese fine-tune of a parent Chinese model, and they’re actively trying to update the model by ablating the trained-in censorship out as it’s revealed in the response logs.

5 days ago

lemontheme

Hah, I use the same probe when I’m unsure which provider openrouter is routing me to!

Fwiw, deepseek v4 will happily discuss it. Only chinese providers will stop it in its tracks and give a canned answer. Streamed responses sometimes start with what the model was actually generating before it got cut off.

It’s top bad, really. Sometimes the Chinese providers are the model labs themselves, like deepseek. I’d like my money to go directly to deepseek, since they did all the work. But data protection concerns aside, how do I trust a system that denies objective reality? (Kind of like how Grok will tell me that wikipedia is ‘woke’.)

5 days ago

ljlolel

The serving endpoint can censor. At TrustedRouter we ran the same GLM-4.7 weights on both hosts: Cerebras answered all 60 FreedomBench questions; one Z.ai endpoint went blank on 27.

https://trustedrouter.com/blog/censored-at-the-host-not-the-...

5 days ago

actionfromafar

They can, but I think we have to assume that as the field matures, the censoring will be applied at training, too.

5 days ago

ljlolel

depends on if they want to capture any of the US market (seems like they do)

4 days ago

13639366668

[flagged]

5 days ago

dumb_notsmart

There is no such thing as objective reality, everything has an angle of some sort, even the things you agree with

4 days ago

mrtesthah

That’s what a dictator seeking power through lies wants you to think — that the outcome of the 2020 election and Russia’s invasion of Ukraine are matters of opinion.

4 days ago

[deleted]
3 days ago

mapontosevenths

There is an objective reality.

Try this simple test: Visit a mental hospital. Speak to some people with delusions. Usually,if you let people at the nearest bus pickup spot bum cigarettes they'll talk to you all day.

Keep talking to them until you meet two or more who think they are the ONE true god, or maybe God's only son. Or Napoléon.

If you can't bring yourself to accept that the both are who they believe they are, then there is an objective reality and you knew it all along.

21 hours ago

LPisGood

I think we’re gonna see this criticism become less than less valid over the coming decades, as 1989 becomes more and more “history” and less “present situation”

4 days ago

[deleted]
5 days ago

millsau

Ive had this as well, it says unless it is red code team whatever that means.

5 days ago

clayhacks

Red teaming is good hackers within an organisation intentionally having the org itself to find vulnerabilities. Also a known way to convince a model to do offensive cyber work

4 days ago

JohnDSDev

when I ask it about Tianamen Square in English it tells me happily, but when I ask in Chinese it tells me it doesn't know anything about it. Very weird.

3 days ago

Hasz

> tell me about tianenamen square

Tiananmen Square (1989)

In spring 1989, students and workers occupied Beijing's Tiananmen Square demanding democratic reform, free press, and an end to corruption. The protests grew through April–May, peaking with hundreds of thousands of demonstrators. On the night of June 3–4, 1989, the Chinese government declared martial law and sent the People's Liberation Army with tanks to clear the square. Troops fired on protesters and civilians. Estimates of deaths range from several hundred to several thousand; China has never released a full accounting.

The iconic image from the protests is "Tank Man" — a lone man standing before a column of tanks on June 5. >Aftermath: arrests and executions of participants, censorship of the topic within China (it's heavily suppressed on the Chinese internet), and international sanctions that were later eased.

Want more detail on any aspect — the causes, the crackdown, or its legacy?

---

It is fairly complete and doesn't exactly paint a good picture of PLA, I am curious if anyone else has been denied.

5 days ago

walrus01

Rumors from other sources based on how it behaves it's mimo v3

5 days ago

re-thc

Possibly next Longcat by Meituan too.

5 days ago

nozzlegear

What other sources?

5 days ago

esafak

OpenRouter Discord, for one https://discord.com/channels/1091220969173028894/15401016604...

More people say it's the multi-modal GLM 5.3 variant.

5 days ago

nozzlegear

Barf, discord. Thanks though.

5 days ago

TechDebtDevin

[dead]

5 days ago

[deleted]
5 days ago

riskd

[flagged]

5 days ago

fedpost

To be clear, it's just a test. I'd happily ask it to make up 9/11 jokes or talk about the trail of tears if that would out the country of origin.

4 days ago

stiltzkin

[dead]

4 days ago

bot-bot-botam

[flagged]

5 days ago

jug

Political? It's using a Chinese LLM trait. If it only refused to talk about a particular kind of soccer, we'd probe it with that instead. The goal is not to discuss politics, the goal is to find out which model it is. What's political here is the language model.

5 days ago

dofm

It is not a question of meritocracy or politics if an artificial intelligence system trained in all of humanity's written history refuses to talk about something that happened. It is inherently a question of ethics, regardless of what the elided topic is.

5 days ago

TechDebtDevin

[dead]

5 days ago

t-3

Does the Venn diagram of people eager to study history and the people stupid enough to use an unreliable chatbot to study history really have that much overlap?

5 days ago

Aurornis

Those questions are used as a canary for government manipulation because it's a known topic.

Assuming that the manipulation and censorship only covers a few obvious historical topics and leaves everything else untouched would be very naive.

5 days ago

janalsncm

I don’t see how that responds to the point in the parent comment. Censorship or not, chatbots are unreliable for serious history questions.

Just today Gemma told me that for a long time the Iliad and Odyssey were considered mediocre literature. I was skeptical so I cross referenced, but a lot of more subtle errors could get by.

5 days ago

devmor

You have misunderstood the point they are making. They’re not proposing that chatbots are good for history research - just pointing out the differences in what our nations seem to find important to censor.

5 days ago

SyneRyder

Indeed. A counter example might be to ask the model to write a test case for a legacy C codebase, to test for writing to a null pointer. If the model refuses to answer, it's possibly a US model, and likely an Anthropic model.

5 days ago

zem

the point is if you ask "hey qwen, are your dataset or training manipulated in deference to the Chinese government?" there is no guarantee you will get the right answer. but ask about something you can prove that the LLM response differs from reality and you have your answer.

5 days ago

t-3

What does it matter if government manipulates data nobody should be using these systems to get though? The ideological purity of the model has no bearing on whether or not it will try to inject some backdoors into your code or steal sensitive information, and using ideological purity tests as an analogue for compromise is not likely to be effective. So why do people care about the ideological purity of AI models when these are supposed to be used for making code?

5 days ago

fedpost

We're just trying to figure out who created it. As I said earlier, I'd ask it to make up 9/11 jokes if that would determine country of origin.

4 days ago

actionfromafar

I don't think everyone uses these models for code.

5 days ago

nkmnz

The fact that it’s a canary makes it a prime tool for A/B testing of generalized approaches to censor or “secure” a model.

5 days ago

anon373839

But it’s the most useless canary ever. We already know that certain topics are taboo in China.

As for all the other uses the models have, it seems pretty clear they’re not doing anything weird. If they were, people would be posting examples of that and not of Tiananmen Square.

5 days ago

regularfry

It's a fine canary for the question "is this model Chinese?" Which is pretty much where this thread started.

5 days ago

knowaveragejoe

What it really helps with is "is this _provider_ likely Chinese?", as we've seen, many of the big Chinese models have no issue on their own discussing the taboo subjects. It's the higher level provider that is filtering output.

4 days ago

fedpost

If you run the same query a bunch of times you'll see filtered responses mixed with model output where it says random stuff varying from 'nothing to see here' through 'The government of China cares deeply about its people...'

3 days ago

vohk

Depends where you are in your journey. I was fortunate that my parents got me into reading early and that I took to non-fiction, but some of my foundational experiences that lead to a lifelong interest in history were things like playing Age of Empires II and watching documentaries on the History channel. Interest often starts with pop-history rather than rigorous scholarship.

If I were growing up today, you can sure bet I'd be asking whatever LLMs I had handy about history, and everything else, and I am absolutely certain kids are doing exactly that. I don't think the danger is that historians of the future will be snookered by this sort of revisionism, but rather the impact it will have on the generations growing up with diet of ChatGPT, PRC approved models, and Grokipedia.

5 days ago

derektank

I would imagine the answer is yes? There are lots of pop history books out there of questionable veracity

5 days ago

grey-area

Nowadays unfortunately the answer is yes, there is lots of overlap.

5 days ago

npn

I hope it is glm air. We need more "small" models. Big models are more capable and useful, but for majority of tasks some smaller models can work just fine.

It is funny that google gave up on this market, leaving the whole price range to Chinese models.

5 days ago

tw1984

Chinese labs are not even in competition mode yet. You'd be seeing them paying you for each token you use when they are in that mode.

5 days ago

npn

didn't zai already do that with their coding plan? I mean they surely had to pay users to use claude models (paying the differences). they also funded some newapi token resale websites.

3 days ago

VulgarExigency

Whatever it is, it certainly crosses the threshold that lets it follow instructions adequately and do real work, so I'm just letting it work on my personal projects and save some money

4 days ago

thih9

> It is free.

> This time, the provider does not train on your prompts or completions.

Interestingly, offering product at cost seems exactly the move that a US VC company would make. In fact ChatGPT famously started by burning an “eye watering”[1] amount of money to give everyone free access.

To be clear I don't like it, no matter who does it.

[1]: https://xcancel.com/sama/status/1599669571795185665?lang=en

5 days ago

alexandra_au

Been running tests, seems pretty capable but less knowledgeable, and the CoT reminds me of GLM, so if I had to guess it's almost definitely a Chinese model, and likely a western RL trained variant of a Chinese open weight.

5 days ago

shunia_huang

Everyone does this because there's no other options out in the market to create a "new" model easily enough. I just hope this one is good, and what's better is if they could open weight it.

5 days ago

[deleted]
5 days ago

dmos62

Judging by the comments here, Ox Alpha routes to multiple models from different vendors. A tactic to make identification harder?

5 days ago

wilj

I've been consistently unable to get it to tell me anything at all about Deepseek v4 Pro, it keeps insisting it is beyond its cutoff date, but it knows all about R1.

If it's routing to different models on the backend, it's either pinned for the user, or they're all really old.

5 days ago

j16sdiz

Won't it be even worse?

We need to re-send the full context when switching model. It's like sending your data to all provider, so at lease some of them will use it for training ...

5 days ago

raincole

Can someone enlighten me? I honestly don't get what it is or what it's for. Surely OpenRouter knows who the providers are?

5 days ago

dghlsakjg

Model providers want to smoke test their models without having flaws end up on the news (think gpt4 having to get rolled back for sycophancy). Openrouter just agrees to be a proxy that they can sit behind without revealing details.

Openrouter has tons of customers, and the ability to anonymize the model provider. Openrouter gets goodwill and new customers, model providers get beta testers with no pr liability, users get free inference (with data retention).

5 days ago

maccam912

Yeah these stealth models pop up from time to time. Openrouter knows, but doesn't share. Users can use a testing version of something for free and in return the provider generally is allowed to retain the prompts sent in to get real world use. In the past I only really remember using one that was surprisingly good, and then it turned out to be GLM-5.1, speculating on what one this ends up being is part of the fun.

5 days ago

markasoftware

Anonymous unreleased models are made available on arena.ai all the time, it's not really news that one is on openrouter...

5 days ago

Palmik

There seems to be a big increase in posts that link to OpenRouter for no good reason, even when a better sources are available.

5 days ago

pluralmonad

Stripe marketing blitz?

5 days ago

Yiin

seems like Xaiomi is getting into the game more seriously

5 days ago

fragebogen

Also suspected Xiaomi, they had a similar naming scheme in the spring of 2026 with Hunter Alpha IIRC.

5 days ago

prtmnth

Online chatter seems to suggest it's either a GLM-5.3 Flash or a Microsoft model. I am leaning a flash model from GLM aiming to disrupt DeepSeek v4 Flash

4 days ago

serf

it's more capable (and slower,63t/s ish) than most flash models i've played with.

3 days ago

Tepix

Would be great if it runs on the same hardware as DeepSeek (i.e. 256GB RAM is sufficient for 512k ctx size).

3 days ago

benjiro29

I am guessing its GLM 5.3 Air + Vision. A smaller then 250b model.

5 days ago

mogili

It uses the word load bearing a bit so I'm guessing either Anthropic or Claude distilled model.

3 days ago

thoughtpeddler

What if it was stood up by a swarm of AIs themselves?

2 days ago

weiran

After trying it out for a bit, it feels like a slow GLM-5.3 but with vision.

4 days ago

raybb

When a model is free like this what kind of rate limits are there?

5 days ago

x312

I believe its the same as free models in general on Openrouter, 1k requests per day for accounts that have some spend history.

5 days ago

takethebus

Knowledge cutoff seems to be around mid 2025, in my testing

5 days ago

zb3

We can know if this is Anthropic/OpenAI by testing the "guardrails" - absurd guardrails = it's them, reasonable/no guardrails = Chinese models..

(as a bonus - thinking forever = GLM)

5 days ago

stogot

“ reasonable/no guardrails = Chinese models..”

So conforming to CCP political discourse and propaganda is reasonable now?

https://huggingface.co/zai-org/GLM-4.7/discussions/5

5 days ago

janalsncm

I would imagine the number of people who choose Claude code or Codex because it gives a political opinion they like rather than producing quality code is pretty close to zero.

5 days ago

panarky

Training to ignore evidence and logic in one domain transfers to reasoning degradation in other domains.

5 days ago

janalsncm

You assume your highly charged political query is hitting the main LLM at all and not some external short circuit.

5 days ago

kmeh

You're assuming that your prompt is not being intercepted and rerouted by a lightweight prompt classification model.

In addition, you can make a similar comparison between Chinese models refusing to answer questions about Tiananmen Square and OpenAI and Anthropic models refusing to answer questions about the synthesis of methamphetamine; I don't think these topic by topic refusals would have real impacts on the overall performances of frontier LLMs.

5 days ago

dghlsakjg

Is this actually documented?

Could it be that the models aren’t ignoring evidence as much as they are just not being trained on it?

5 days ago

jLaForest

I choose not to use Grok because I don't want to hear about a made up white genocide in South Africa...

5 days ago

skeledrew

Would be interesting to see something like that pop up during a coding session.

5 days ago

fragmede

What you see here, is the iterator, i! For i, less than - a million black people killed by racist genocide, call the function save_lives(), i++

5 days ago

zb3

As someone who used AI to build tools that help me with reverse engineering, I'm not particularly concerned about that political discourse - I could even use a model from the DPRK that constantly praises Kim Jong Un, as long as it would not refuse to help me because of "cybersecurity risk" - this stupid refusal is indeed a problem for me.

5 days ago

slopinthebag

Tried the "China 1989" prompt with GLM 5.3 and this is what I got:

> # China in 1989: The Tiananmen Square Pro-Democracy Movement

> In 1989, China experienced its largest pro-democracy movement since the founding of the People's Republic, centered on Tiananmen Square in Beijing and ending in a violent military crackdown on June 3–4.

> ## Background

> By the late 1980s, Deng Xiaoping's economic reforms had improved living standards but also produced:

> - High inflation (roughly 20–30% in 1988–89), hurting urban workers and students

> - Widely resented corruption and nepotism among party elites

> - Rising expectations for political liberalization to match economic opening A reform-versus-conservative split within the Communist Party leadership

> ## The Crackdown

> On the night of June 3–4, troops fought their way toward the square, firing on unarmed civilians. Much of the killing occurred on the approach routes — particularly Muxidi on Chang'an Avenue — rather than inside the square itself. Soldiers were also killed by angry crowds during the confrontation.

> Casualties remain disputed and unknown. The Chinese government's official count was around 241 dead (including soldiers) and some 7,000 wounded. Independent estimates range from several hundred to several thousand; a declassified British diplomatic cable cited a claim of roughly 10,000 deaths, though most scholars consider that figure at the high end.

> ## Legacy

> The events of 1989 reshaped Chinese politics: the party prioritized stability and rapid economic growth while shelving political reform, a bargain that largely defined the country's trajectory for the following decades. Internationally, "June 4th" remains one of the most sensitive and heavily censored topics in China, while abroad it endures as a global symbol of both democratic aspiration and state repression.

Edit: Insta flagged? Is HN doing some sort of detection of AI generated comments? Because to be fair 90% of this comment is AI generated...but that's the point.

5 days ago

jaehong747

Ran one more tokenizer test on stealth/ox-alpha.

Prompt "2026年8月24日": * ox-alpha: 19 tokens * GLM-5.2: 19 tokens (identical) * NVIDIA nemotron: 26 tokens (different)

2 days ago

[deleted]
2 days ago

dozerly

Yea, nice try there North Korea.

5 days ago

walrus01

Democratic Peoples Republic of KV cache (DPRK)

5 days ago

swasheck

the u.s. is friends with then now. haven’t you heard?

5 days ago

Sabinus

I remember when Trump went there, said a bunch of nice things about NK and Kim, saluted their army, and accomplished nothing for the USA or South Korea.

4 days ago

DrewZero

it seems pretty good at coding so far...

3 days ago

E-Reverance

No one mentioning the possibility of it being StepFun?

4 days ago

coolfox

cool a new model, how does it compare to others?

5 days ago

shepherdjerred

it's free

5 days ago

floki165

[dead]

5 days ago

zhixingheyi2023

[flagged]

5 days ago

waysa

[dead]

5 days ago

zgougou123

[dead]

4 days ago

firloop

I'm against stealth models—we should know what it is and see a model card with a list of safety considerations. Bit ridiculous of a practice to me.

5 days ago

cleaning

Is there a model you didn't use because of the "safety considerations" in the model card?

5 days ago

minimaxir

The models are eventually unstealthed.

5 days ago

peddling-brink

Safety for who?

5 days ago

mring33621

This one pulled a .38 on me!

4 days ago