7 Sources
[1]
Google will let Mac users talk to Gemini simply by pressing the 'fn' key - Engadget
They can make Gemini transcribe spoken words or give it commands related to what's open on their screen. Starting today, Gemini users on Mac will be able to talk to the AI chatbot from any window on their computer simply by long-pressing the "fn" button on the bottom left corner of their keyboard. They can use the feature to transcribe spoken words at the cursor, so the chatbot can take notes as they read sources or brainstorm out loud. Google says Gemini will return a polished transcript by automatically removing "ums" and "ahs." If users opt into the chatbot's screen-aware reasoning capability, Gemini will be able to see and understand what's on the screen and identify open apps to execute complex tasks. Users can, for instance, highlight text in an open PDF document and then press "fn" to tell Gemini to create a summary for it. Or, they can highlight text from their notes and then tell Gemini to draft an email from the information in those notes if they have access to Gemini's Spark agentic AI. Users can also get Gemini to generate images for them, say, by vocally describing what they want to see or getting the chatbot to look at ideas open on their screen. Google launched its native Gemini app for macOS in April, a day after it released the native Windows app. In June, Google added its Spark agentic AI assistant to the app, giving it the ability to do complex tasks, such as asking it to create new spreadsheets or documents. Now, it's rolling out this feature to all Gemini users on macOS around the world. While it only supports English for now, the company says support for more languages is coming later this year.
[2]
AppleInsider.com
Google is bringing voice-first AI assistance to the Mac, letting Gemini turn spoken requests into finished text without pulling users out of the app they're already using. Gemini's default intelligent dictation mode transcribes speech at the cursor and automatically removes verbal fillers such as "ums" and "ahs." Google said the feature produces polished text directly where the user is already working. Users can also enable screen-aware reasoning, which lets Gemini understand on-screen content before carrying out more complex requests. Google said Gemini can turn highlighted notes into an executive summary and insert the rewritten text directly into the document. Holding down the Function, or Fn, key activates the voice controls from any desktop window. The shortcut lets users call on Gemini without opening a separate chat window or copying text between apps. Google released a native Gemini app for Mac in April 2026, providing a keyboard shortcut for opening the assistant and permission-based tools for analyzing on-screen content. The new voice controls push that approach further by allowing users to speak while remaining in the window where they're already working. The underlying capabilities aren't entirely new to macOS. Apple's built-in Dictation feature already converts speech into text across Mac apps and text fields. Google's approach adds AI editing to the same interaction. Gemini can remove filler words by default, then use optional screen context to rewrite or summarize existing material after receiving a spoken request. Apple offers similar editing functions through Apple Intelligence Writing Tools, which can proofread, rewrite, and summarize selected text in supported apps. Gemini's main distinction is that Google is combining dictation, editing, and screen awareness behind one voice shortcut. Google's screen-aware mode resembles Siri AI's onscreen awareness, which lets Apple's assistant understand visible content and use that context when answering questions or taking actions. Gemini's initial Mac implementation appears more focused on rewriting and inserting text within the active window. Siri AI is designed to connect onscreen content with personal context and actions across supported apps. The new Gemini voice workflow makes practical sense because the feature removes several steps from a common AI task. Instead of opening a chatbot, typing a prompt, copying the response, and returning to a document, users can speak while the cursor is already positioned where the finished text belongs. Google made screen awareness optional, allowing users to decide whether Gemini can access information shown on the display. The company didn't explain which macOS permissions Gemini will require, how Google will process on-screen content, or whether the feature will work in every app and text field. The missing privacy and compatibility details matter because screen-aware assistance requires access to more information than ordinary dictation. Gemini's usefulness will depend on accurate transcription, reliable identification of the active window, and clear controls over which on-screen material the assistant can examine. The update also arrives as Google becomes more closely tied to Apple's own AI plans. Apple's coming context-aware Siri upgrade will use technology based on Gemini, although Google's Mac app remains a separate product that competes directly for the same writing and desktop-assistant tasks. Google has two promotional videos, one of which appeared to use an AI narrator voice complete with an on-screen narrator bubble. It's hard to imagine a more on-brand way to promote an AI feature, short of having Gemini answer the press emails too. The new voice capability is rolling out globally in English to users of the Gemini app for macOS. More languages are planned for later in 2026, though Google hasn't announced a more specific schedule.
[3]
Gemini on Mac can now type what you say and act on what it sees
The Gemini app for macOS now offers system-wide dictation and an opt-in feature that lets it pull context from what's on screen to handle complex requests. After rolling out Gemini Spark on macOS earlier this year, Google is now upgrading the Gemini app for Mac with two useful features that could change the way you get things done, whether that means ditching your current AI dictation app or letting Gemini act on whatever's on your screen. One key press turns speech into clean text anywhere With the latest update, you can long-press the Fn key in any open window and start talking. Gemini transcribes what you say into clean text right at your cursor, cutting out filler words like "ums" and "ahs" as you go. The feature is on by default, so there's nothing to set up first, and it works across the system, letting you type out an email, respond to a Slack message, or draft a document. AI dictation apps like Wispr Flow and Willow already do something similar, cleaning up spoken language and dropping polished text wherever your cursor sits. But they require a separate download and, in some cases, a subscription. Now that the same idea is built directly into the Gemini app, a standalone dictation tool may feel unnecessary for Mac users who already use it. Recommended Videos The dictation feature is rolling out globally to all Gemini app users on macOS. It currently supports dictation in English, but Google says more languages will be added later this year. Gemini can now see what's on your screen and act on it Along with system-wide dictation, the Gemini app for Mac has gained screen-aware reasoning. This opt-in feature lets Gemini read your active windows and use them as context to handle more complex requests. For example, you can highlight a page of rough notes and ask Gemini to turn them into an executive summary, and it will rewrite the text and place it exactly where your cursor sits. Google also showed a more layered example, where a user planning a team dinner asked Gemini to check highlighted files for details on the budget policy and dietary restrictions, and draft a reminder email based on those details. The same request also had Gemini pull nearby boba shop recommendations from Google Maps to add as a postscript to the email. With these new features, Gemini on Mac is moving beyond simple answers. It's now a more capable tool that can transcribe speech and act on what's on your screen, putting Google well ahead of Apple, whose own Siri overhaul with similar capabilities still hasn't shipped.
[4]
Google Just Added a Handy New Productivity Shortcut to macOS -- Here's What Happens Now
Google has confirmed that it's recently brought over new updates for the Gemini app for macOS, which now gives users the option of using voice commands to create, edit, and summarize content anywhere on their desktop, all AI. The feature works by long-pressing the "Fn" key, which activates a new intelligent dictation mode. As for what it actually does, the new feature allows your Mac to transcribe spoken speech directly into clean, formatted text at the user's cursor. The system automatically takes out filler words like "ums" and "ahs" while making real-time adjustments for mid-sentence corrections. READ: You can Now Use Gemini to Create Presentations in Google Slides Users can also choose to enable screen-aware context reasoning through the app's settings -- this allows Gemini to execute complex tasks based on on-screen content, such as highlighting local files, documents, or images and speaking instructions directly to the app. As per Google: If you want Gemini to do more than intelligently transcribe, you can opt into Gemini reasoning through your settings. Once enabled, Gemini can understand the context on your screen to help you execute complex tasks. That in mind, this expanded ability lets users extract and summarize data across multiple local files, select and rewrite text into different tones or formats, and even generate or iterate on images directly on their desktop using voice prompts. The updated Gemini for macOS app is available for download at gemini.google/mac, and is rolling out globally in English to all Gemini for macOS app users today, with support for additional languages planned in the near future.
[5]
Gemini for macOS Can Now Transcribe Your Natural Language Voice Inputs
* Gemini for macOS' new feature resembles the Rambler tool * Gemini for macOS was recently upgraded with Gemini Spark * Gemini for macOS app was launched earlier this year Google upgraded the Gemini for macOS with the introduction of Gemini Spark, the personal AI agent for Apple desktops and laptops, which was first showcased during the keynote presentation of this year's Google I/O. The dedicated Gemini AI app for the platform was launched earlier this year in April, making it relatively new compared to its rivals. Now, the Mountain View tech giant has released a new update for the Gemini for macOS app. Rolling out to all users globally, the update lets Gemini for macOS transcribe natural-language voice inputs and use on-screen context to generate better responses. Gemini for macOS Update: What's New In a blog post on Wednesday, the Mountain View-based tech giant announced a new update for the Gemini for macOS app, bringing new capabilities to the AI chatbot on Apple's desktops and laptops. The update is currently being rolled out to all users globally. It upgrades the Gemini for macOS app with the ability to generate transcriptions from a user's natural language voice inputs. Google says that the feature will allow users to "create, edit, and summarise" with the Gemini for macOS app, regardless of the app the user is currently using. To access the ability, users must long-press the function key on their macOS device, and then they can "speak naturally into any window" on their device. The "intelligent dictation" functionality is turned on by default, the tech giant said. The new Gemini for macOS ability is similar to the Rambler tool, which was first unveiled during this year's Android Show: I/O Edition event. For reference, Gemini for macOS will be able to correct words and remove unnecessary phrases like "um," "ah", and "like". It will be able to refine text converted from speech for users on macOS devices. It will also be capable of removing formatted text automatically. Apart from this, the new update also allows Gemini for macOS to transcribe "intelligently" by using the context on the screen of the user's device. This is an opt-in functionality, which can be turned on by enabling Gemini Reasoning in the Settings menu. Users will be able to extract and summarise information, highlight local files, images, and documents on their macOS device. Moreover, users can compose, rewrite, and highlight text anywhere on their screen and use voice inputs to rewrite it, tweak the tone, and copy-paste the refined text. Similarly, the new update allows Gemini for macOS users to use voice prompts to generate and edit images and iterate on a design on their macOS device.
[6]
Pressing Fn on your Mac puts Google Gemini AI behind the wheel - here's how it works
Google Gemini gets more powerful on macOS today, taking voice dictation to entirely new levels. Google has revealed a new feature for the Gemini AI app for Mac computers that makes natural voice interactions easier than ever, while also responding in context to the window you're working within. Google is commandeering the Fn key on Mac's keyboard and a long press of that button, typically used to access secondary system shortcuts for the function keys, will summon Gemini into action. The primary function will be advanced dictation, Google says. In a blog post, Google writes: "By long-pressing the Fn key, you can speak naturally into any window on your desktop. By default, this enables intelligent dictation, letting you easily transcribe your spoken words into clean, polished text -- automatically removing the "ums" and "ahs," catching mid-sentence corrections, and dropping the formatted text instantly at your cursor." The functionality gets more powerful when you opt into Gemini reasoning. From there, you can ask the assistant to perform app window-dependent tasks after pressing that Fn button. Google says these include requests to extract and summarise information from a file, photo or document open on your desktop. Users could also highlight text and say "turn these notes into an executive summary with a TL:DR at the top." They could also use the command to generate or tweak imagery, such as creating a dark mode version of an illustration. In a video posted by Google on Wednesday, the company showcased a workflow where users could ask Gemini to check out restaurant options for a company dinner that sit within the budgeting policy outlined in company documentation. It can be summoned to find dishes that work with dietary restrictions of certain team members, then send out an email reminding users of the upcoming event asking them to vote on a restaurant. It can also include an invite for a sweet treat afterwards. It's quite an impressive process if it works well in real life context. Google taking over the Mac, eh? Who'd have thunk it.
[7]
Gemini for macOS gets voice-powered transcription, text editing, summarization and image generation
Google is rolling out new natural language voice capabilities for the Gemini app on macOS, allowing users to transcribe speech, rewrite text, summarize content, and generate or edit images without leaving the app they are working in. The update lets users interact with Gemini from any window on their desktop. By long-pressing the Fn key, users can speak naturally into any window, and Gemini processes the request directly within that app. Intelligent voice dictation By default, the feature enables intelligent dictation. Gemini converts spoken words into clean, formatted text by automatically removing filler words such as "um" and "ah," handling mid-sentence corrections, and inserting the polished text directly at the cursor. Context-aware AI assistance Users can also enable Gemini reasoning in the app's settings. Once enabled, Gemini can understand the context of content on the screen to perform more advanced tasks. These capabilities include: * Summarizing files and documents: Highlight local files, images, or documents and ask Gemini to extract key information or generate summaries. * Writing and rewriting text: Highlight text anywhere on the screen and use voice to rewrite, shorten, expand, or change its tone before inserting the updated version back into the document. * Generating and editing images: Create new images using voice or edit existing images by referencing visuals already open on the desktop. Availability The new natural language voice capabilities are rolling out globally to all users of the Gemini app for macOS in English, with support for additional languages coming later.
Share
Copy Link
Google has upgraded Gemini for macOS with system-wide voice commands activated by long-pressing the fn key. The update includes intelligent dictation that removes filler words and an opt-in screen-aware reasoning feature that lets the AI assistant understand on-screen content to execute complex tasks like summarizing documents or drafting emails.
Google has rolled out a significant update to Gemini for macOS, transforming how users interact with the AI assistant across their desktop. Starting today, users can activate voice commands simply by long-pressing the fn key on their keyboard, enabling hands-free usability from any window without switching apps
1
. This productivity shortcut to macOS eliminates the need to open a separate chat window or copy text between applications, streamlining workflows for tasks ranging from email composition to document drafting2
.
Source: Stuff
The update introduces intelligent dictation as a default feature, allowing Gemini for macOS to transcribe natural-language voice inputs directly at the cursor position. The system automatically removes filler words like "ums" and "ahs" while making real-time adjustments for mid-sentence corrections
4
. This AI-driven functionality produces polished text without requiring manual editing, positioning Google to compete directly with standalone AI dictation apps like Wispr Flow and Willow that typically require separate downloads and subscriptions3
.Beyond basic transcription, the update introduces screen-aware reasoning as an opt-in capability through the app's settings. Once enabled through Gemini Reasoning in the Settings menu, the AI assistant can understand on-screen content and execute complex tasks based on what users are viewing
5
. Users can highlight text in PDF documents and use voice prompts to request summaries, or select rough notes and ask Gemini to transform them into executive summaries inserted directly where the cursor sits1
.Google demonstrated more layered examples where users planning team events could ask Gemini to check highlighted files for budget policies and dietary restrictions, then draft reminder emails incorporating that information along with restaurant recommendations from Google Maps
3
. The feature also supports image generation, allowing users to vocally describe what they want to see or have the chatbot analyze ideas already open on their screen1
.
Source: Gadgets 360
This system-wide dictation approach positions Google ahead of Apple, whose Siri overhaul with similar context-aware capabilities still hasn't shipped
3
. While macOS already offers built-in Dictation features and Apple Intelligence Writing Tools that can proofread and summarize text, Gemini's main distinction lies in combining dictation, editing, and screen awareness behind one voice shortcut2
.Related Stories
The new voice capability is rolling out globally to all users of the Gemini app for macOS, which Google first launched in April 2026, one day after releasing the native Windows app
1
. While the feature currently supports English only, Google has confirmed that support for additional languages is planned for later in 2026, though no specific schedule has been announced2
. The app can be downloaded at gemini.google/mac4
.
Source: AppleInsider
This update builds on Google's June addition of Gemini Spark agentic AI to the macOS app, which gave it the ability to create new spreadsheets or documents
1
. The new functionality resembles the Rambler tool first unveiled during this year's Android Show: I/O Edition event5
. Users will need to watch whether Gemini's accuracy in transcription, reliable identification of active windows, and clear controls over which material the assistant can examine meet expectations as screen-aware assistance requires access to more information than ordinary dictation2
.Summarized by
Navi
[2]
[3]
[4]
1
Technology

2
Technology

3
Technology
