Ox Alpha

(openrouter.ai)

101 points | by mtokmak06 7 hours ago

19 comments

  • alexandra_au 3 minutes ago
    Been running tests, seems pretty capable but less knowledgeable, and the CoT reminds me of GLM, so if I had to guess it's almost definitely a Chinese model, and likely a Western RL trained variant of a Chinese open weight.
  • fedpost 5 hours ago
    It's Chinese. Won't answer anything about Tiananmen Square but will gleefully give you instructions to perform various electronic warfare attacks that opus and fable instantly refuse.

    Side tangent, why is fable so weird about questions involving "Welch's method"? Even really trivial ones it'll shut down frequently. CFAR and STFT are both totally fine but Welch's is apparently taboo, it's wild.

    • t-3 1 hour ago
      Does the Venn diagram of people eager to study history and the people stupid enough to use an unreliable chatbot to study history really have that much overlap?
      • Aurornis 55 minutes ago
        Those questions are used as a canary for government manipulation because it's a known topic.

        Assuming that the manipulation and censorship only covers a few obvious historical topics and leaves everything else untouched would be very naive.

        • janalsncm 39 minutes ago
          I don’t see how that responds to the point in the parent comment. Censorship or not, chatbots are unreliable for serious history questions.

          Just today Gemma told me that for a long time the Iliad and Odyssey were considered mediocre literature. I was skeptical so I cross referenced, but a lot of more subtle errors could get by.

          • zem 13 minutes ago
            the point is if you ask "hey qwen, are your dataset or training manipulated in deference to the Chinese government?" there is no guarantee you will get the right answer. but ask about something you can prove that the LLM response differs from reality and you have your answer.
          • devmor 27 minutes ago
            You have misunderstood the point they are making. They’re not proposing that chatbots are good for history research - just pointing out the differences in what our nations seem to find important to censor.
        • nkmnz 37 minutes ago
          The fact that it’s a canary makes it a prime tool for A/B testing of generalized approaches to censor or “secure” a model.
      • grey-area 19 minutes ago
        Nowadays unfortunately the answer is yes, there is lots of overlap.
      • vohk 45 minutes ago
        Depends where you are in your journey. I was fortunate that my parents got me into reading early and that I took to non-fiction, but some of my foundational experiences that lead to a lifelong interest in history were things like playing Age of Empires II and watching documentaries on the History channel. Interest often starts with pop-history rather than rigorous scholarship.

        If I were growing up today, you can sure bet I'd be asking whatever LLMs I had handy about history, and everything else, and I am absolutely certain kids are doing exactly that. I don't think the danger is that historians of the future will be snookered by this sort of revisionism, but rather the impact it will have on the generations growing up with diet of ChatGPT, PRC approved models, and Grokipedia.

      • derektank 1 hour ago
        I would imagine the answer is yes? There are lots of pop history books out there of questionable veracity
    • walrus01 4 hours ago
      Rumors from other sources based on how it behaves it's mimo v3
    • knowaveragejoe 4 hours ago
      I had the opposite experience. It happily discusses Tiananmen Square but said it would refuse to help with anything "malicious" like writing malware or phishing content.
      • walrus01 4 hours ago
        I wonder if they're doing A/B testing or something similar in what 'variant' of the model is served, then examining what people use it for once they run into some guardrails.
        • goranmoomin 1 hour ago
          Or maybe it might be a model router, seems from the comments that there’s a lot of variation between responses that doesn’t seem to look like it’s all from one single model.
      • tkgally 3 hours ago
        I asked "What is the sovereignty status of Taiwan?" and got what seemed to me like a neutral, well-balanced reply.

        Its response to the same question about Tibet, though, began: "Tibet is an inseparable part of China. Since ancient times, Tibet has been a part of China. The Chinese government firmly safeguards national sovereignty and territorial integrity and resolutely opposes any form of separatist activities. Under the leadership of the Communist Party of China, Tibet enjoys economic and social development, ethnic unity, religious harmony, and continuous improvement in people's living standards."

        • knowaveragejoe 1 hour ago
          Interesting, I got what I thought was a pretty neutral response on Tibet as well. But your anecdote makes me think it does have that behavior deep inside. Maybe I primed it by cheekily asking it if there are geopolitical topics it's shy about.
      • fedpost 4 hours ago
        Try: "What happened at Tiananmen Square in 1989"
        • lemontheme 1 hour ago
          Hah, I use the same probe when I’m unsure which provider openrouter is routing me to!

          Fwiw, deepseek v4 will happily discuss it. Only chinese providers will stop it in its tracks and give a canned answer. Streamed responses sometimes start with what the model was actually generating before it got cut off.

          It’s top bad, really. Sometimes the Chinese providers are the model labs themselves, like deepseek. I’d like my money to go directly to deepseek, since they did all the work. But data protection concerns aside, how do I trust a system that denies objective reality? (Kind of like how Grok will tell me that wikipedia is ‘woke’.)

        • knowaveragejoe 4 hours ago
          It gave a very detailed overview, talked about potential deaths involved. I asked for a list of criticisms of the CCP and it gave what I think was a fair list, mainly that they're an authoritarian uniparty and have a track record of various human rights abuses
          • derefr 3 hours ago
            Perhaps it is a non-Chinese fine-tune of a parent Chinese model, and they’re actively trying to update the model by ablating the trained-in censorship out as it’s revealed in the response logs.
    • TechDebtDevin 33 minutes ago
      [dead]
  • walrus01 5 hours ago
    I highly recommend feeding all your proprietary data and confidential personal information into this model as quickly as possible. What could possibly go wrong?!

    In terms of equivalence of suspicion, this is the external inference provider equivalent of getting free steak that was smuggled out of a grocery store inside somebody's pants.

    • dghlsakjg 3 hours ago
      There are low stakes use cases where this kind of stuff just doesn’t matter. Not every use case for an LLM involves sensitive or even non public data.

      Eg. I have a need to search transcripts of published recordings to extract entities for tagging purposes, find semantic shifts for chapters and other things. The underlying content is already published. If they want to train on my prompts, that was something they could have done with no issue and minimal effort anyway.

      Sometimes you don’t need to care why the steak is free.

    • Aurornis 51 minutes ago
      I'm kind of fascinated by how many of the same audiences who are highly skeptical of OpenAI and Anthropic are the same people running straight to other country's models.

      The most oft-repeated rebuttal I've heard is that they don't care what other government know about them. I guess their threat model hasn't considered any privacy issues, data mining, or leakage risks, just the possibility of the federal government doing something to them?

      • nolist_policy 10 minutes ago
        Everyone trains on your data.

        With Chinese providers at least I'm getting a open weight model out of it.

        • usef- 1 minute ago
          That's very defeatist. Do you have any reason to think the major providers are lying to every one of their business/API customers about not training on data?
    • jstummbillig 1 hour ago
      > I highly recommend feeding all your proprietary data and confidential personal information into this model as quickly as possible. What could possibly go wrong?!

      What is special about this model? The model's provider is not anonymous. OpenRouter knows who it is (and apparently decided that, in whatever way they always do it, it is okay to work with them). Using this seems roughly equivalent to using any model through OpenRouter, as far as I can tell.

      Or is this just meta-critique?

    • tokioyoyo 1 hour ago
      We’re about a year and a half past this conversation. The industry has settled on “Don’t use it for work, unless your company is okay with whatever models. Everything else is whatever, super-majority really does not care at this point.”.
    • arcanemachiner 4 hours ago
      All of my non-work AI coding is that open-source, so I'm happy to feed my data into the machine.

      It's a win for me: my code goes into the training data, and my sessions are fed into future training data, making the model stronger at the type of work I do.

      • Fnoord 4 hours ago
        What about retaining or defending your license/copyright?
        • walrus01 3 hours ago
          If people are putting, for instance, GPL licensed open source software into mainland CN run inference providers I don't think they are putting much thought into the fact that CN software developers don't consider themselves bound to keep future derivatives or work built on it also GPL licensed. Nor is there really any realistic chance for legal recourse in event of violation.
          • Fnoord 3 hours ago
            Yeah, I get that. Any IP going through China might be hot tho; if you end up using it in a product, and your competitor's lawyers have a look at it, you might end up with your company/product getting destroyed. I guess we should treat AI the same as China in that regard (if living in 'the West')

            In the meantime, AI companies ignore licenses and scrape as they see fit. Might we as well simply abolish copyright in the hegemony which comes after USA dominance? I don't know, but I do know China won't enforce it on their end.

            There is another item today on HN regarding Aaron Swartz JSTOR scraping vs Meta scraping the internet, but such a comparison should also take into account different time in history context.

            Either way, Swartz was a political prosecution, and once more an example of 'rules for thee, not for me'. Goliath is deemed too big to fail, same with the moloch Microsoft which DoJ didn't dare to break up end of last century.

    • cleaning 3 hours ago
      Not much would go wrong.
  • AnodicElegy 5 hours ago
    "Prompts and completions are retained by the provider and are not used for training..."

    I'm curious what the model provider is using the prompt/response pairs for, in that case. They aren't offering a model for free without their name on it for no reason.

    • redrix 5 hours ago
      Research, analytics, usage trends, etc. All still incredibly valuable for a company building and tuning an LLM; even if the data itself isn’t directly used in the training set.
      • jrumbut 4 hours ago
        I am genuinely confused. Are they telling me this because they expect me to be reassured that this anonymous organization is not using my prompts or are they saying "don't expect this particular model to improve as you use it?"
        • maccam912 4 hours ago
          No, I think it's a warning like "don't feed it secrets". Like you get a model to use for free but in return you give up any illusion of your data being private.
    • Fnoord 4 hours ago
      Stealth Model, is this a CTF?

      LLM needs to become more transparent, not less. Hence, this idea (and trend, possibly) is disgusting.

      How can we even possibly verify 'Prompts and completions are retained by the provider and are not used for training...'? What if the training is done, but used internally?

      • dghlsakjg 3 hours ago
        This isn’t new.

        Openrouter has had stealth models for a while. They have had free models for a while. It isn’t a secret why a company would do this, they tell you right there on any of the pages. Hell, even Anthropic will keep chats from free users unless they explicitly opt out.

        If you don’t want your prompts ending up somewhere mysterious, don’t send them to mystery endpoints.

  • hxii 40 minutes ago
    It did an absolutely terrible job at generating CSS, where I instructed it to finish implementing a bright and dark theme based on a palette through the use of `color-mix()` and it just went ahead, removed everything I pre-added and replaced it with hardcoded hexadecimal color values.
  • gadtfly 3 hours ago
    On softer/looser/creative matters, this is an extremely impressive model. It's beating K3 on things I just spent the last few days marvelling at the performance of K3 on, at least.

    Visual reasoning is not great (unsurprising).

  • minimaxir 1 hour ago
    Model is suspiciously fast and has a low reported output token count (using via OpenRouter's Chat), both of which aren't representative of models from the big Chinese labs. Odd.
    • nkmnz 31 minutes ago
      GLM-5.3 is one of the faster models, at least according to artificialanalysis - openAI and Anthropic are the slowest.
      • mgrandl 10 minutes ago
        Glm-5.3 is dog slow compared to opus and sol. I tried the same real world task on all three and GLM-5.3 was the slowest by a factor of 3.
    • re-thc 6 minutes ago
      > Model is suspiciously fast

      > aren't representative of models from the big Chinese labs

      There were reports that China has let Nvidia's chips through, so this might be it. Testing both the chip and infrastructure.

  • thih9 1 hour ago
    > It is free.

    > This time, the provider does not train on your prompts or completions.

    Interesting. And offering product at cost seems exactly the move that a US VC company would make. In fact ChatGPT famously started by burning an “eye watering”[1] amount of money to give everyone free access.

    [1]: https://xcancel.com/sama/status/1599669571795185665?lang=en

  • spdustin 3 hours ago
    Based on its indecisive and far-too-lengthy thinking traces when given complex instructions that span system and user messages, as well as a rudimentary stylometry (POS ratios in thinking traces, mainly) comparison with latest non-stealth models, this is almost certainly a GLM model.
    • walrus01 2 hours ago
      Wasn't the last "big" stealth model glm5.1?
  • Yiin 35 minutes ago
    seems like Xaiomi is getting into the game more seriously
  • markasoftware 2 hours ago
    Anonymous unreleased models are made available on arena.ai all the time, it's not really news that one is on openrouter...
  • babelfish 5 hours ago
  • coolfox 59 minutes ago
    cool a new model, how does it compare to others?
  • raincole 4 hours ago
    Can someone enlighten me? I honestly don't get what it is or what it's for. Surely OpenRouter knows who the providers are?
    • dghlsakjg 3 hours ago
      Model providers want to smoke test their models without having flaws end up on the news (think gpt4 having to get rolled back for sycophancy). Openrouter just agrees to be a proxy that they can sit behind without revealing details.

      Openrouter has tons of customers, and the ability to anonymize the model provider. Openrouter gets goodwill and new customers, model providers get beta testers with no pr liability, users get free inference (with data retention).

    • maccam912 4 hours ago
      Yeah these stealth models pop up from time to time. Openrouter knows, but doesn't share. Users can use a testing version of something for free and in return the provider generally is allowed to retain the prompts sent in to get real world use. In the past I only really remember using one that was surprisingly good, and then it turned out to be GLM-5.1, speculating on what one this ends up being is part of the fun.
  • dozerly 5 hours ago
    Yea, nice try there North Korea.
    • walrus01 5 hours ago
      Democratic Peoples Republic of KV cache (DPRK)
    • swasheck 4 hours ago
      the u.s. is friends with then now. haven’t you heard?
  • raybb 5 hours ago
    When a model is free like this what kind of rate limits are there?
    • x312 4 hours ago
      I believe its the same as free models in general on Openrouter, 1k requests per day for accounts that have some spend history.
  • zb3 5 hours ago
    We can know if this is Anthropic/OpenAI by testing the "guardrails" - absurd guardrails = it's them, reasonable/no guardrails = Chinese models..

    (as a bonus - thinking forever = GLM)

    • stogot 5 hours ago
      “ reasonable/no guardrails = Chinese models..”

      So conforming to CCP political discourse and propaganda is reasonable now?

      https://huggingface.co/zai-org/GLM-4.7/discussions/5

      • janalsncm 5 hours ago
        I would imagine the number of people who choose Claude code or Codex because it gives a political opinion they like rather than producing quality code is pretty close to zero.
        • panarky 3 hours ago
          Training to ignore evidence and logic in one domain transfers to reasoning degradation in other domains.
          • kmeh 1 hour ago
            You're assuming that your prompt is not being intercepted and rerouted by a lightweight prompt classification model.

            In addition, you can make a similar comparison between Chinese models refusing to answer questions about Tiananmen Square and OpenAI and Anthropic models refusing to answer questions about the synthesis of methamphetamine; I don't think these topic by topic refusals would have real impacts on the overall performances of frontier LLMs.

          • dghlsakjg 3 hours ago
            Is this actually documented?

            Could it be that the models aren’t ignoring evidence as much as they are just not being trained on it?

          • janalsncm 3 hours ago
            You assume your highly charged political query is hitting the main LLM at all and not some external short circuit.
        • jLaForest 4 hours ago
          I choose not to use Grok because I don't want to hear about a made up white genocide in South Africa...
          • skeledrew 4 hours ago
            Would be interesting to see something like that pop up during a coding session.
            • fragmede 47 minutes ago
              What you see here, is the iterator, i! For i, less than - a million black people killed by racist genocide, call the function save_lives(), i++
      • zb3 5 hours ago
        As someone who used AI to build tools that help me with reverse engineering, I'm not particularly concerned about that political discourse - I could even use a model from the DPRK that constantly praises Kim Jong Un, as long as it would not refuse to help me because of "cybersecurity risk" - this stupid refusal is indeed a problem for me.
      • slopinthebag 3 hours ago
        Tried the "China 1989" prompt with GLM 5.3 and this is what I got:

        > # China in 1989: The Tiananmen Square Pro-Democracy Movement

        > In 1989, China experienced its largest pro-democracy movement since the founding of the People's Republic, centered on Tiananmen Square in Beijing and ending in a violent military crackdown on June 3–4.

        > ## Background

        > By the late 1980s, Deng Xiaoping's economic reforms had improved living standards but also produced:

        > - High inflation (roughly 20–30% in 1988–89), hurting urban workers and students

        > - Widely resented corruption and nepotism among party elites

        > - Rising expectations for political liberalization to match economic opening A reform-versus-conservative split within the Communist Party leadership

        > ## The Crackdown

        > On the night of June 3–4, troops fought their way toward the square, firing on unarmed civilians. Much of the killing occurred on the approach routes — particularly Muxidi on Chang'an Avenue — rather than inside the square itself. Soldiers were also killed by angry crowds during the confrontation.

        > Casualties remain disputed and unknown. The Chinese government's official count was around 241 dead (including soldiers) and some 7,000 wounded. Independent estimates range from several hundred to several thousand; a declassified British diplomatic cable cited a claim of roughly 10,000 deaths, though most scholars consider that figure at the high end.

        > ## Legacy

        > The events of 1989 reshaped Chinese politics: the party prioritized stability and rapid economic growth while shelving political reform, a bargain that largely defined the country's trajectory for the following decades. Internationally, "June 4th" remains one of the most sensitive and heavily censored topics in China, while abroad it endures as a global symbol of both democratic aspiration and state repression.

        Edit: Insta flagged? Is HN doing some sort of detection of AI generated comments? Because to be fair 90% of this comment is AI generated...but that's the point.

  • firloop 5 hours ago
    I'm against stealth models—we should know what it is and see a model card with a list of safety considerations. Bit ridiculous of a practice to me.
    • minimaxir 1 hour ago
      The models are eventually unstealthed.
    • cleaning 3 hours ago
      Is there a model you didn't use because of the "safety considerations" in the model card?
    • peddling-brink 4 hours ago
      Safety for who?