Google and OpenAI moved on the same day to the "Fast AI" route, but the rhythm of their landings is different. Google is officially on line Gemini 3.7 Flash and is open to developers; OpenAI has previewed GPT-5.6 Sol Ultrafast, a high-speed model that is currently only invited for testing to selected clients.
Gemini 3.7 Flash is fully open.
Google has positioned Gemini 3.7 Flash as a generic model for programming and AI proxy scenarios. The model supports text, image, video, audio and PDF input, and can also call on tools and control computers with the goal of dealing with multistep tasks and reducing manual intervention.
According to the data disclosed in Google, Gemini 3.7 Flash can receive up to 1 million inputs token, with an output ceiling of 64,000 token. According to Google, the version performed better than previous generations on software engineering, web development and complex knowledge tasks, while responding faster.
Google, for example, in its tested coding mission, Gemini 3.7 Flash spent 2 minutes 13 seconds, while the previous generation of Flash took more than 5 minutes. At the same time, Google emphasized that the new model was not intended to reduce quality in exchange for speed.
The price continues to go down.
In addition to speed, Google uses low-cost as a selling point. According to the information in the text, Gemini 3.7 Flash was priced by the end of this year:
- Token charge $ 0.75 per million
- Token per million
- This price is about half of Gemini 3.6 Flash's initial price.
The above-mentioned preferential prices will be increased to $1.50 per million of inputs and $7.50 per million of exports after 31 December.
In terms of performance comparison, Google cited the home-based benchmarking test that Gemini 3.7 Flash was leading 11 out of 18 test categories, including the Code Arena web development score of 1588 Elo and 30.4 per cent of the Enterprise Workstream Testing Bench. However, these results are based on Google's own methods.
OpenAI main hit faster.
OpenAI is not a new model but a Ultrafast model for GPT-5.6 Sol. According to the report, this model is supported by Cerebras, which can run at about 14 times the GPT-5.6 Sol standard.
Cerebras' crystal-circle chip is said to produce 750 tokens per second, approximately 560 English words per second. OpenAI seeks to highlight the significance of this speed for real-time scenes such as voice agents, for example, through faster completion of reasoning and response during phone calls.
However, unlike the direct opening of Google, GPT-5.6 Sol Ultrafast is currently available only to a small number of clients in OpenAI API, and will be gradually opened up to a larger number of business users as the economy expands. OpenAI has also not published a positive systematic comparison of the high-speed model with other models, with more references to client feedback at this stage.
The focus of competition shifted to real-time agents.
This simultaneous release reflects the changing focus of competition for AI manufacturers. The market focus is moving from a mere comparison of “who is smarter” to “who is fast enough to support real-time agents”. Speed is becoming a more direct indicator of experience for products requiring continuous assignments and immediate response to users.
Google is a pioneer in terms of current availability. Gemini 3.7 Flash has been online in over 160 countries and territories and can be called directly by developers; the hypervelocity model of OpenAI is still at the invitation stage, with limited short-term coverage.
