Last Updated: 12 December 2024

For over two decades, Google’s main goal has been to organize the world’s information and make it easy to use. The first Gemini model, launched last December, could understand text, images, audio, and code all at once. This helped many developers and made Google’s tools more useful.
Now, Gemini 2.0 takes a big step forward. Instead of just showing you information, it can think ahead and perform tasks for you.
These “agentic” abilities let it plan several steps, understand more complex requests, and use online tools on your behalf—all while you stay in control. In other words, Gemini 2.0 is built to not only understand the world but also to help you act on it.
Before, Gemini focused on understanding what you asked for. With Gemini 2.0, it can also figure out what to do next. This shift is backed by years of research and special hardware called Trillium TPUs. With these tools, Gemini 2.0 can work with multiple types of information, from videos and images to audio files, and provide answers in many forms.
Starting today, some developers and trusted testers can try out Gemini 2.0’s Flash experimental model and Deep Research feature.
Deep Research breaks down complex topics, searches for the best details, and shares clear reports.
The idea is to make AI help you more directly, whether you’re solving a problem at work or exploring a new subject at home.
Search has always been a key way to find information online. AI already changed how we search, making it easier to ask new kinds of questions and understand more data.
With Gemini 2.0’s improved reasoning, these AI-powered search overviews will soon handle even tougher questions, solve complicated math problems, understand images, and help with coding tasks.
These features are still in testing, but they’ll come out more widely next year. The goal isn’t just to give you a quick fact—it’s to help you truly understand a topic so you can decide what you think about it.
Gemini 2.0 Flash is a faster, more powerful version of Gemini that can handle images, video, audio, and text. It can even create images and speak back to you with text-to-speech. Plus, it can use tools like Google Search or code execution without you having to do everything yourself.
Developers can access these features through the Gemini API, Google AI Studio, and Vertex AI. Basic features are open to everyone, while more advanced ones, like image generation or TTS, will be tested with a small group first. The new Multimodal Live API lets AI work with live audio and video streams, offering a more interactive way to use AI in real time.

Source: Google
With Gemini 2.0, “agentic” AI can do more than just answer questions. It can understand long instructions, plan tasks, and take action online. This opens the door to new experiences:
Project Astra is a test version of a universal AI assistant. Thanks to Gemini 2.0, it can now speak multiple languages, use Google tools like Search and Maps, remember details from past chats, and respond faster. Some testers are even trying Astra on smart glasses, exploring how AI might fit into everyday life.
Project Mariner, introduced Wednesday, shows how an AI agent can browse the web on your behalf. It takes screenshots of your browser, understands what’s on the page, and can do tasks like adding items to a shopping cart. It’s still slow and limited for safety reasons, but it hints at a future where you say what you need, and the AI does the clicking.
Jules helps developers with coding tasks directly in GitHub. It can read code, fix issues, and make changes on its own, while a human developer oversees it. This could simplify coding and speed up the development process.
Gemini 2.0 doesn’t just work with websites. It can also help you navigate video games by giving suggestions based on what’s happening on the screen. Google DeepMind is testing this with game companies like Supercell. The same reasoning skills could one day help AI agents understand and assist in real-world settings, like robotics.
When AI can take action on your behalf, safety and trust become more important than ever. Google is working with experts and using Gemini 2.0’s reasoning to find and fix problems before they happen. Users have control over their data and can easily delete sessions, and the AI must follow certain rules—like asking permission before doing sensitive tasks.
As Gemini 2.0 improves, the way we find and use information may change. Instead of clicking through pages, you might guide an AI agent to do it for you. This approach could save time, open up new possibilities, and help people make better decisions. It’s the start of a new era in how we work with information.