New Gemini 3.5 Flash Models Are Faster and Cheaper but Not Smarter

New Gemini 3.5 Flash Models Are Faster and Cheaper but Not Smarter

Google introduced three new Gemini models on Tuesday, showing the vendor is paying attention to cost concerns in the AI model market while still competing against OpenAI and Anthropic by making its models faster and more domain-specific.The Gemini 3.6 Flash model is coding and knowledge focused. With this model, Google concentrated on cost efficiency, with a lower cost per output tokens on some benchmarks compared to Gemini 3.5 Flash. Flash-Lite is the most cost-effective model in the series, delivering about 350 output tokens per second, Google said. And 3.6 Flash Cyber is a new cybersecurity model that will be paired with CodeMender, Google’s AI agent for detecting, patching and preventing security vulnerabilities.The new models arrive before the Gemini 3.5 Pro model, which Google is currently testing and will be made available soon. However, as updated Flash series arrives as momentum in the worldwide AI market shifts toward less expensive models, with Chinese AI vendors Alibaba and Moonshot this week launching new high-performing, low-cost models.Related:OpenAI Urges Enterprises to Use Its Scorecard to Measure Worth of AI“Cost utilization of AI models in general in the industry has become a topic of interest,” said Sid Nag, president and chief research officer of AI analyst firm Tekonyx. “Cost optimized with speed is going to be a continuing dialogue that we as industry participants are going to be talking about.”The Focus on CostOne of the biggest enterprise costs is model training and inference. So, if Google can reduce latency and token cost with 3.6 Flash, it can “translate into substantial savings for large deployments, especially enterprise deployments,” Nag continued.Compared to recent models, Gemini 3.5 Flash-Lite appears to be the cheapest option in the line, costing $0.3 per million input tokens and $2.5 per million output tokens. By comparison, the Chinese model Kimi K3 costs $3 per million input tokens and $15 per million output tokens. Gemini 3.5 Flash-Lite is also cheaper than OpenAI’s GPT-5.6 models, which range from $1 to $5 per million input tokens and $6 to $30 per million output tokens.However, Google did not completely hit the mark on cost with all the new models it released, said Ayham Boucher, the Head of AI Innovations and a lecturer in information science at Cornell University. The cost of 3.6 Flash is still higher per unit of output than that of models from SpaceXAI, for example. Gemini 3.6 Flash runs at $1.50 per million input tokens and $7.50 per million output tokens. Comparatively, SpaceXAI’s Grok 4.5 costs $2 per million input tokens and $6 per million output tokens.Related:Alibaba Qwen 3.8 Max Shows China Closing in on U.S. Models“As we all know, with the input tokens with agentic systems, your prompt could be very short, and then the agent is going to go and generate a lot of output before its response,” Boucher said. “Output tokens do matter more than input tokens.”Speed vs. IntelligenceBoucher noted that, in addition to price, Google optimized the speed of its models.“They definitely achieved the speed, which is definitely way ahead of all the frontier closed models, and you could only compete with it with open source models,” Boucher said.However, regarding intelligence, he said that Google itself failed to reach the level of other closed models from Anthropic and OpenAI. This does not mean that Google is falling behind, but rather that it is targeting infrastructure to enable faster, cheaper models, he said.“Cost of inference, and managing high throughput of inference, is going to be very pivotal and essential as we continue to deploy AI systems successfully,” Boucher said. “This is what Google is betting on, and they are not wrong that there is a huge market where you do not necessarily need the smartest model. You need a fast, smart model that is cheap.”He added that if the Gemini 3.6 version expected next year is more intelligent than the current Fable version from Anthropic, while also being fast and less expensive, it could garner a lot of interest. This is because next year’s Fable will likely be more intelligent and many enterprises do not always need the most intelligent model because many tasks that don’t require the highest intelligence.Related:Startup Launches Foundry to Develop Materials for ChipsHowever, Google’s focus on more affordable models does not mean it should not also provide a frontier model that also competes on intelligence, Boucher said.“If it doesn't, then it has fallen behind in the race of frontier intelligence,” he said. “Just because you're making your models faster, you shouldn’t also give up your spot in the leading edge of intelligence.””The Cyber ModelIt is also clear that Google is focusing on cybersecurity models with its release of Gemini 3.5 Cyber. For Nag, the cyber model illustrates another way AI is becoming more specialized.“It has three major enterprise implications,” he said. “AI becomes part of the software development lifecycle. AI becomes a force multiplier for security teams, and of course, lower-cost security agents become practical.”While with Gemini 3.5-Cyber, Google has created an agentic system that autonomously identifies and fixes security problems, there is still much to learn about this model, said TJ Marlin, CEO of Guardrail Technologies.“If you're going to have an autonomous system that's finding the vulnerabilities and fixing them, where are the checks and balances there?” Marlin said. “There’s a number of known limitations, reliability issues and risks that don’t go away just because they created an orchestration model.”In addition, while Google has showcased the positives of Cyber, it is unclear how reliable the model is, he said.“We're not learning the full picture,” Marlins said. “All the risks that exist for AI being manipulated intentionally or unintentionally, resulting in bad outcomes, they also exist for this agentic system that Google has created for vulnerability detection.”About the AuthorNews Writer, AI BusinessEsther Shittu has covered AI technologies and industry trends since 2021. As co-host of the Targeting AI podcast, she talks with experts, thought leaders and practitioners exploring critical AI developments. Before AI Business, she wrote for SearchEnterpriseAI, the New York Daily News, Bklyner and the Brooklyn Daily Eagle. When she's not diving deep into the world of AI, she spends her time on passion projects and raising her three daughters.

Original Source

Read the full article at Aibusiness →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.