From Wednesday, October 9, 2026, free Gemini is being reduced to a single model: according to Google's updated official help page, users without a subscription in the app will only have Flash-Lite, while Flash and Pro will become unavailable. Meanwhile, the AI Pro plan at $19.99 per month will get the Deep Think mode for the first time, previously available only to AI Ultra subscribers, and in October the app will introduce low, medium, and high effort levels, which will consume the compute limit faster.

image
image
image

What happened

Google updated the official help page 'Changes to Gemini model access and limits' (support.google.com/gemini/answer/17004136): starting Wednesday, October 9, 2026, in the Gemini app, accounts without a subscription will only get the Flash-Lite model — according to 9to5Google, specifically Gemini 3.5 Flash-Lite. The Flash (3.6) and Pro (3.1) models are disappearing from the free tier. Subscribers to the AI Plus plan at $4.99 per month will keep Flash-Lite and Flash, and Google promises to notify the cutoff date for this plan by email. The AI Pro plan at $19.99 per month and AI Ultra retain all three models, with AI Pro opening the Deep Think mode for the first time, described as a maximum parallel reasoning mode that until now was only available on Ultra ($99.99 and $199.99 per month). The update also affects the interface: in October, the app will introduce low, medium, and high effort levels, which improve task performance but consume the compute limit faster.

Context

This is a continuation of the long restructuring of Gemini monetization. In May 2026, Google tied app usage to compute limits that update every 5 hours up to a weekly cap, and the boundary between plans was based on quota volume. Now the boundary becomes the model class itself: the free tier is left with the smallest and lightest model in the lineup, while heavy and 'smart' computations are moved to paid plans. This direction aligns with competitors' decisions and Google's own releases: in September, Anthropic left the Claude Opus 5.5 model only for paid users, and the new Gemini 4 Argon model was also initially given to AI Ultra subscribers. A separate trace of the change will be left on model comparisons: the word 'Gemini' means different models and different compute budgets depending on the plan, so measurements before and after October 9 or between plans are only correct when the model version, plan, and effort level are fixed.

Why this matters for the industry

For the industry, the key is that Google is for the first time splitting Gemini plans by model class, not just by quota volume: the lightest model remains free, while access to Flash, Pro, and Deep Think is monetized in a range from $4.99 to $19.99 per month. Flagship inference is effectively leaving the free acquisition funnel: demos, freemium prototypes, and workflows built on the free app lose their support, even though the model can be removed in a matter of days. Product teams should pay attention to the new interface pattern: low, medium, and high effort levels make 'quality at the expense of compute' an explicit configurable option, and the visible quota becomes part of the design. The drop of Deep Think from Ultra to AI Pro allows two readings: either the cost of parallel reasoning has decreased, or Google is differentiating the mid-tier plan — there is no measurable quality increase in the sources. Practically, companies have a short window until October 9 to recalculate scenarios dependent on free Pro, and if interested, launch capture campaigns targeting the audience leaving the free tier.

Why this matters for users

Users should do a quick review of tasks tied to Pro in free Gemini: long documents, code, and math are better finished before Wednesday, October 9, after which the model will become unavailable, and the free tier context will be limited to approximately 32,000 tokens. The remaining limit is visible on gemini.google.com in the Settings → Usage Limits section — it shows what part of the quota has already been used. Then there are three options: AI Plus at $4.99 per month with Flash-Lite and Flash models, AI Pro at $19.99 per month with all three models and Deep Think, or switching to an alternative chatbot. For everyday requests like chats, translations, and short queries, the difference from Flash-Lite will probably be unnoticeable; it appears on large documents, complex code, and math. With the high effort level after its introduction, caution will be needed: it improves performance quality but noticeably eats up the compute limit faster.

What is still unknown / limitations

The conclusion that the cost of flagship inference is not covered by the free funnel is not supported by the sources: there is no data on inference prices, unit economics, or cost structure, so this is a hypothesis about Google's motive, not a fact. It is unknown how the low, medium, and high effort levels will be deducted from the limit; there are no benchmarks of their effect in the sources either. Deep Think lacks a measurable quality increase. No information has been published about changes to the Gemini API: token prices, latencies, and quotas are probably not affected yet, but this should be checked separately. The model cutoff dates for AI Plus have not been announced — Google promises to name them by email. The forecast that by early 2027 the norm will be a free light model plus a flagship by subscription is an interpretation based on Anthropic and Google's decisions, not a statement.

Sources

Author

Look at AI, editorial team