Google Gemini is back with powerful free AI upgrades

07 Aug 2026 04:07 11,156 views
Google has quietly rolled out a wave of Gemini and NotebookLM updates that turn its AI into a true everyday assistant. From voice-to-any-app dictation and agent workflows to short‑form video, music generation, and free app hosting, here’s what’s new and how you can use it.

After a few quiet months, Google has come back strong with a wave of Gemini and NotebookLM updates that focus less on benchmarks and more on real-world usefulness. Instead of chasing the top spot on leaderboards, Google is wiring AI directly into your desktop, browser, documents, and even app hosting—much of it for free.

Talk to any window: Gemini’s new “Speak to window” dictation

One of the most impressive new features is a system-wide voice assistant that can listen to you and act inside whatever window you’re using. It’s like dictation on steroids: it doesn’t just transcribe your voice, it understands the full context of the app or page you have open.

On the Gemini desktop app (currently on macOS, with Windows coming soon), you’ll see a new microphone icon. You can also assign a global keyboard shortcut in the settings under “Speak to window.” Once set up, you can hold the key and start speaking in any app—Docs, email, a blank note, or even a browser form—and Gemini will instantly type what you say.

Unlike classic dictation tools, this works in multiple languages and can use on-screen context. For example, with an email open, you can say something like: “Write a polite reply in Spanish thanking them for the update in this email,” and Gemini will read the email above, understand it, and generate a full response instead of just transcribing your words.

It also tracks context across windows. You can, for instance, open a tweet about a new AI model in one window, ask Gemini to summarize it as an email for your team, then switch back to Gmail and let it draft the message based on what it saw on X. You can select existing text and ask Gemini to rewrite it, translate it, or expand it, and it will operate directly on the selected content.

All of this is available for free in the desktop app, making it a powerful everyday assistant for writing, replying, and organizing without constant copy-paste.

Gemini Spark: Google’s answer to AI agents

Google is also rolling out a new agentic mode called Gemini Spark. Think of it as Gemini’s version of AI coworkers like Claude Workbench or GPTs with tools: instead of simple chat, you get an AI agent that can plan, execute multi-step tasks, and connect to external services.

From the Gemini web interface, you’ll find a Spark tab (though it’s still restricted in some regions, like the EU, where you may need a VPN to see it). Inside Spark, you can create “tasks” that the agent will handle for you, such as managing your Gmail, Google Calendar, Google Drive, and other parts of the Google ecosystem.

The interface is built around workflows rather than single prompts. When you ask what Spark can do, it shows a breakdown of capabilities: email management, scheduling, research, automation, and more. It can chain steps, ask for confirmations, and act as a proper assistant rather than a simple text generator.

Connecting third‑party apps via MCP and Zapier

One of the biggest changes is that Spark is no longer limited to Google tools. It now supports “Connected applications” via MCP (Model Context Protocol), which lets you plug in external services like Zapier.

To connect a tool like Zapier, you paste its MCP URL (you can usually find this in the app’s documentation or by asking Gemini itself), confirm the permissions, and authorize the connection. Once linked, Spark can see all the actions that Zapier exposes—potentially thousands of apps and automations.

From there, you can use the @ mention system inside tasks (for example, @Zapier) and ask Gemini to create automations such as: “Create a workflow that turns new rows in my spreadsheet into Instagram posts.” Spark will set up the Zap, show you the steps, and guide you through final authentication. This finally breaks Gemini out of the purely Google ecosystem and into your full tool stack.

Skills and scheduling: reusable workflows

Gemini Spark also introduces “skills”—reusable mini-workflows that define how the agent should handle a certain type of task. You can describe a skill in natural language (for example, “Analyze my best-performing content each week and propose 10 new ideas”) and let Gemini build it, or you can create and import skills manually.

Skills can be scheduled to run automatically at specific days and times, and they can operate across both Google and third‑party tools. Because they run in the cloud, your computer doesn’t need to be on. This turns Gemini into a true automation hub for recurring tasks like content analysis, lead generation, or social posting.

For readers interested in the broader context of Google’s agent push, it’s worth also checking how this fits into the company’s wider enterprise agent strategy in Google’s Gemini Enterprise agent platform.

NotebookLM: from research hub to short‑form video generator

NotebookLM, Google’s AI notebook tool, has also received some meaningful upgrades that make it more useful for research, learning, and content creation.

Collections for organizing notebooks

You can now group multiple notebooks into “collections.” When you click “Create collection,” NotebookLM shows all your existing notebooks so you can bundle them into themed sets—like “University notes,” “Marketing research,” or “Tokyo travel project.”

This is especially helpful if you use NotebookLM as a knowledge base: each collection becomes a focused workspace where you can chat with multiple related notebooks at once.

AI‑generated Shorts from your notes

The standout new feature is a built‑in short‑form video generator. Once your Google account language is set to US English, a new “Short” option appears under the “Video overview” section of a notebook.

From there, you can ask NotebookLM to create a vertical video (similar to YouTube Shorts, TikTok, or Instagram Reels) based on the notebook’s content. For example, you can request: “Create a short video about the top 10 secret places in Tokyo.” NotebookLM then generates a 1–2 minute video with narration, images, and even simple animations that match the story.

You can preview the video, listen to the AI voiceover, and download it for sharing on social media. This works especially well for educational content, language learning, or storytelling, as NotebookLM can turn structured notes into ready-to-post short videos.

New Gemini models: 3.6 Flash and 3.5 Flashlight

Google has also launched new Gemini models: Gemini 3.6 Flash and Gemini 3.5 Flashlight. They’re not designed to dominate intelligence benchmarks, but to balance quality with speed and cost so they can power real products at scale.

On a capability vs. cost chart, these models sit in the middle: not as “smart” (or expensive) as top-tier frontier models, but efficient enough to run across Google’s ecosystem—Gemini chat, NotebookLM, creative tools, and more.

Gemini 3.6 Flash: interactive, visual, and versatile

Gemini 3.6 Flash is a general-purpose model that shines in interactive and visual tasks. In Canvas mode, you can upload something like a hand-drawn floor plan from an architect and ask Gemini to turn it into an interactive app for clients.

From a rough sketch, Gemini 3.6 Flash can generate a full interface where you can:

• Inspect and edit room layouts and dimensions
• Simulate walking paths and lighting based on orientation and time of day
• Add and move furniture, with warnings when items overlap walls or doors
• See a running budget for all selected items and work
• Compare the original plan with an optimized proposal side by side
• Export a detailed PDF report for clients

This kind of rapid, interactive prototype is especially useful in fields like architecture and real estate, where visualizing options and costs quickly can save a lot of time.

Gemini 3.5 Flashlight: ultra‑fast analysis and dashboards

Gemini 3.5 Flashlight is optimized for speed and large-scale analytical tasks. It may not handle complex visual layouts as well as 3.6 Flash, but it excels at processing lots of structured or semi-structured data.

In Canvas, you can upload a batch of invoices and receipts and ask it to:

• Extract key fields (company, date, currency, base amount, tax, etc.)
• Aggregate and analyze spending
• Build an interactive dashboard with charts and summaries

The model parses the documents, writes the code for the dashboard, and renders it in seconds. You can then ask for additional charts or breakdowns and see them update almost instantly. This makes Flashlight a strong choice for finance, operations, or any workflow where you need quick, repeated analysis over many files.

Lyria 3.5: Google’s new music generation model

Another quiet but important release is Lyria 3.5, a music generation model that’s already among the best options for AI-composed tracks with vocals.

You can access it through compatible tools like Flow Music. In the settings, select the Lyria 3.5 model, describe the song you want (style, mood, instruments, language, etc.), and the tool will generate full lyrics plus multiple musical variations.

Once you pick a version you like, you can go further and ask it to generate a music video concept or even a simple video clip. You can specify portrait or landscape format, and the system will adapt the visuals accordingly. The result can then be downloaded for refinement or direct sharing.

Creative video tools: Flow and Bits

Google’s creative stack is also evolving, especially around video.

Flow: community tools and custom pipelines

Within Flow, which is designed for building longer AI-generated videos, there’s now a “Tools” section with templates from both Google and the community. These tools can:

• Change the visual style of footage
• Add vector graphics or overlays
• Apply text effects and transitions
• Chain multiple effects into reusable pipelines

You can use existing tools or build your own, combining models and effects to create custom workflows for your specific content style.

Bits: AI‑powered video creation and editing

On the Bits platform, you can now create AI videos from text, from a set of images, or using avatars. There’s support for custom avatars (with some limits on using real faces), and the underlying video model is currently ranked among the best for AI video generation.

Bits also lets you edit existing videos using AI. You can, for example, change a character’s clothing, alter the background scene, or restyle the entire clip in a different aesthetic, all while preserving motion and structure. This makes it a powerful tool for content creators who want to iterate quickly on visuals without reshooting footage.

Free app publishing with Google AI Studio

Finally, Google AI Studio now makes it easy to host and share AI apps for free, even if they were originally built elsewhere.

In AI Studio, you can create a new app and use the “Import from GitHub” option. Once you connect your GitHub account, you can select any repository—say, a real estate assistant app you built with another model—and AI Studio will deploy it into Google’s environment.

After migration, the app runs inside AI Studio with full access to its interface. You can scroll through the UI, test features, and then enhance it by connecting Gemini models and Google services from the settings panel.

When you’re ready to share, go to the “Publish” section. You can choose a custom subdomain (for example, yourappname.studio). As long as the name is available, AI Studio will reserve it for you and publish the app at that URL, which anyone can access in their browser. The first person to claim a name owns it, so short or brandable names are worth grabbing early.

This turns AI Studio into a no-cost launchpad for prototypes, internal tools, or even public-facing apps, tightly integrated with Gemini and the rest of Google’s ecosystem. For a deeper dive into how Google is positioning this platform for developers and businesses, you can look at this practical guide to the Gemini Enterprise Agent Platform.

Why these updates matter

Individually, each of these features—voice dictation, agents, short-form video, music, app hosting—might not seem revolutionary. Together, they show a clear strategy: embed AI deeply into everyday workflows rather than just chasing benchmark scores.

With Gemini now able to listen to any window, automate tasks across apps, turn notes into videos, analyze documents at scale, compose music, and host full applications, Google is building an AI layer that quietly sits under everything you do. The real impact won’t come from any single announcement, but from how you plug these pieces into your own projects and daily work.

Share:

Comments

No comments yet. Be the first to share your thoughts!

More in Gemini