On August 20th, an anonymous AI model called Ox Alpha appeared on OpenRouter, OpenCode, Cline and Nous Research platforms and was made available free of charge. In the absence of the company's signature, this “invisible model” quickly aroused concern in the development circles.
The heat rise was initially driven by a set of coded baseline results that were rapidly disseminated. The developer Ben Davis, who tested in the 10 task samples of DeepSWE, gave a first-round pass rate of 80%, up from 65% of Claude Fable 5 and 52% of GPT-5.6 Sol. This result was subsequently transmitted in large numbers, giving the outside world the impression that Ox Alpha was clearly ahead of the existing head model.
Complete test returned about 63%.
DeepSWE is a baseline test for code-oriented agents to measure whether or not the first attempt will solve the real GitHub problem. The test consisted of 113 jobs and covered 91 code warehouses.
On the basis of the two separate and complete run points mentioned in the paper, all of the results of Ox Alpha fell below about 63 per cent, roughly at the same level as GPT-5.6 Sol, rather than stabilizing higher than the latter. In other words, the "80% lead" that was previously widely disseminated came more from small sample results and more volatile space.
As of 22 August, Ox Alpha was still not on the list of mainstream models, such as Artificial Analysis or LMSys Arena, which also means that public evaluation data on it are still limited.
Free and open for attention
In addition to the performance discussions, the open conditions offered by the platform have led to increased attention. OpenCode states that it can support 10 trillion tokens per day; Nos Research says that its own Portal capacity can reach 100 trillion tokens per day.
- Free for about a week.
- Support 1 million context windows
- Support multi-mode input and zero data retention
Stripe CEO Patrick Collison also wrote on platform X that the model was “very impressive”. This short evaluation further magnifies the discussion of Ox Alpha.
Fingerprint analysis points to the brain.
There are still no public institutions to claim Ox Alpha. Multiple rounds of comparison have been developed around their sources.
It was mentioned that some common candidates had been partially excluded. For example, Mimo v2.5 Mimu supports audio input, while Ox Alpha expressly rejects audio; DeepSeek has not been able to launch video in the past and usually issues open source weights instead of an anonymous preview; Google and Qwen do not match Ox Alpha in tokenizer and video encoding.
The rest points more to AI. Fingerprint tests by independent researchers show that the tokenizer of Ox Alpha is fully consistent with the GLM-5.3 in the standardized tests, that its video token consumption is also fully compatible with the GLM-5V-Turbo in four video samples, and that it displays error features similar to the GLM series and rejection of audio input.
However, there is still a doubt about this judgement. GLM-5.3 was still a pure text model when it was released on August 14, while Ox Alpha had full video capability when it went online on August 20. Based on this time difference, if the above-mentioned fingerprinting is established, Ox Alpha is more likely to be an undisclosed multi-modular upgrade of the GLM-5.3 series than a direct reproduction of existing models.
The term “invisible model” usually means an unsigned development agency that throws a model through a route-by-line platform such as OpenRouter to collect real-use data before making it official. It is mentioned that this practice has become more common in some of China ' s AI laboratories. According to OpenCode, the free opening period for Ox Alpha lasted about 27 August, when its real identity was disclosed and remained one of the market concerns.
