Audio input interests me when it preserves the request a person finished making. In its June 3, 2026 announcement, Google presented Gemma 4 12B as a notebook option with native audio input. That deserves a place in a comparison because the work begins before the answer: somebody has to get the material into the model.
Later that month, on June 20, 2026, I asked for an assessment of the week's experiences and a comparison between models. I wanted to look back at the work and understand the options. Revisiting this announcement adds a requirement to that comparison: examine how the request arrives. Counting parameters while ignoring input leaves a practical part of the job out of the calculation.
Take a hypothetical interface request. Someone begins by saying to remove a button, reconsiders, and asks for it to remain visible but disabled. The expected result changed halfway through the explanation. A beautifully written task asking for removal is still wrong. The agent needs to recognize the correction instead of turning both versions into simultaneous requirements.
That case can be investigated through transcription followed by a model, or through direct audio input. I want to compare what each route preserves against the original recording. Accepting the file is only the beginning. The response has to represent the final intention, including hesitation and changes of mind. People speak that way; they do not deliver a reviewed specification with every breath.
The notebook adds another concrete requirement. Before building a workflow around this, I need a configuration capable of running the task. Memory and response time belong in that decision. Google's proposal puts the model among the candidates; the announcement alone neither selects a configuration nor shows how it behaves on the work I want done.
I am also unwilling to manufacture audio for a request that already exists in writing. Converting input just to use a new feature creates more work. The capability makes sense when an explanation begins as speech and the existing route loses something or takes unnecessary effort to produce a task.
My evaluation starts with the correction example: supply the recording and check that the final task asks to disable the button. Then compare input routes and the cost of execution in the chosen environment. If the corrected intention disappears, native audio failed on the exact detail that made me interested in it.