DeepReinforce released Open Source Model Series Ornith-1.0 last week and is open for MIT permission at Hugging Face. The series does not focus on general dialogue, but is oriented towards an AI-coded intelligence body capable of performing its tasks independently in a code warehouse and terminal environment.

Provide 9B to 397B versions

The model was released with four specifications: 9B, 31B, 35B MoE and 397B MoE. DeepReinforce defines it as a model family for a smart body-coding mission, with a focus on having the model perform multiple steps in a real development process, rather than generating a single code.

Intelligent coding refers to the ability to read files, run tests, position errors, change codes and continue to recycle until the task is completed. Unlike common chat AI, these models place more emphasis on completing the full development process with less manual intervention.

Training methods to cover implementation strategies

DeepReinforce describes how many coding smarts usually rely on artificially predefined implementation frameworks, such as when to call the tool, how to break down the task, and how to deal with the error. Orith ' s approach is to incorporate this implementation framework into training so that the model can gradually optimize its implementation strategy while fulfilling its mandate.

At the same time, the company mentioned that three layers of restriction were added to the training to avoid a model obtaining high marks through a drill rule: The test environment and the validation script are non-modifiable; the control system identifies access to restricted paths or changes to the certification process; and a freeze-free award model is used to reject abnormal results.

397B and 9B both give high marks.

According to DeepReinforce, the 397B flagship score was 82.4 on SWE-bench Verified. This test gives the model a real open-source warehouse a loophole or flaw, allowing it to complete the repair without seeing the test set and score it in proportion to the success of the solution.

  • 397B score on SWE-bench Verified 82.4
  • 397B score on Terminal Bench 2.1 77.5
  • 9B score on SWE-bench Verified 69.4

According to the report, the 397B version of SWE-bench Verified is 80.8 higher than Claude Opus 4.7 and 80.6 higher than DeepSeek-V4-Pro; in Terminal Bench 2.1, Ornith is 77.5 and Claude Opus 4.7 is 70.3.

However, there have been training pollution disputes around SWE-bench, i.e. models may remember the answer to the questions they have seen in training. To this end, Ornith has also published results on the more difficult SWE-bench Pro, with the 397B version divided into 62.2, which is lower than its performance on Verified but remains competitive.

What is more interesting is probably small model performance. The 9B version reached 69.4 on SWE-bench Verified, up from 52.0 in Gemma 4-31B and close to 70 partitions in Qwen 3.5-35B. In terms of the size of the parameters, this means that Orith's small model displays greater efficiency in coding tasks.

Unequivocal general chat model

DeepReinforce also reminds in the model document that Ornith-1.0 may not be suitable for non-coding tasks. The model may not be a suitable option if the user wishes it to summarize the document, mail or process the general text.

According to the article, Ornith is well positioned: It serves teams that have been set up to host coding water lines, smart infrastructure or automated development tools, rather than general AI assistants for ordinary consumers.

By contrast, Ornith's performance over and above the partial closed-source model is mainly reflected in the specific coding benchmark and is more appropriately understood in the context of the same open-source model and intelligently coded mission. For developers, the real value lies in whether smaller and medium-sized models can operate steadily in marginal hardware or in a self-build environment.