量子位

No more waiting for Gemini 4! Google has released a work agent that supports calling Claude.

A new “seam monster” has appeared, how can it remain stagnant?

Hengyu, from Ao Fei Temple
QuantumBit | Public Account QbitAI

Good good good, Google has alsojoined the battle of office AI agents, and by emphasizing multi-model invocation, it even put its competitor Claude on the list!

This morning, Google Cloud released a long article titled “Welcome to Gemini at Work 2026: Introducing the Gemini agent,” which was signed by Google Cloud CEO Thomas Kurian. The article highlights a series of updates to Gemini in the field of enterprise Agents.

Among them, the most significant is the newly released Gemini Agent.

According to Google’s vision, this will be a general office agent that covers various types of work tasks and can operate long-term.

It can help you search for information, write emails, create PPTs, analyze data, write and run code, call tools across applications, plan tasks on your own, and gather multiple sub-Agents together when necessary.

And there are two things that are particularly worth noting.

First, it can have its own corporate identity, becoming an AI colleague with an email address, a calendar, and a separate account.

Second, it can automatically select the underlying model based on the task, and even call Claude from Anthropic.

Wow, even Google might not necessarily use its own Gemini?

What’s even more interesting is that all the features released this time seem to be aimed at the recently booming office Agent market.

Xiao Za would probably have the following internal thoughts:

Song, don’t spin up that stupid Gemini model of yours, just turn your head and come at me?????

Gemini agent

The working principles that Google has set for it are those of Aunt Chili:

You give it objectives, not instructions.

Give it a goal, rather than step-by-step instructions.

That is, I want the Gemini agent to help with the work, without having to tell it what to do first and what to do next step by step.

Just report the final desired result, and let it arrange the rest by itself.

Its core capabilities include a unified entry point, persistent operation in the cloud, cross-platform calls, multi-Agent collaboration, event triggering, and scheduled tasks.

The example mentioned in the text is to ask it to help complete a market analysis report.

It can search and organize information on its own, analyze relevant data, create financial models in Google Sheets, and then use Google Slides to generate a presentation.

Throughout the entire process, it can select tools and call different applications as needed for the task, and finally return the completed work to you.

You might have a lot of questions in your mind—Similar cross-app operations aren’t anything new, so what are you doing?

Most netizens are also not optimistic about this. Below the official announcement tweet, there is a sketch in the black-pink style:

But this time, Google’s release of Gemini Agent emphasizes the effort to consolidate various working capabilities into a single Agent.

What’s different?

Under the large framework of the Gemini Agent, Google has also introduced the Coworker Agent, which literally means “colleague agent”.

Google now allows companies to create a long-term, clearly defined role for Gemini, enabling it to participate in daily collaboration like other members of the team.

Moreover, the AI colleague has a separate Google Workspace account, corporate email, calendar, and Google Drive storage space… and can also appear in the company address book.

The initial barrier to getting started is essentially none. After creation, team members can invite it into the Google Chat group like they would with any other colleague, or @ it in a document.

The AI colleague uses their own identity; when performing operations, they leave records of themselves, which facilitates enterprise management and tracking of what exactly they do.

At this point, it still looks pretty normal..

What differs is that, in addition to its independent identity, the Gemini Agent has a more complete memory mechanism.

According to the official introduction, it has a total of four types of memory.

The first is session memory.

Primarily used to save the context of the current task.

Even if a task lasts for several days, the Gemini Agent can remember which steps have been completed and what remains to be done next.

The second type is semantic memory (Semantic Memory).

The Gemini Agent accumulates knowledge from the documents it reads, interactions with people, and collaborations with other Agents, and organizes it into a structured knowledge base.

For example, remember which products the company has, what each team member is responsible for, and what a specific business term means within the company.

The third is procedural memory (Procedural Memory).

Primarily used to record how a task should be completed.

For example, what data is required for a certain type of report, what analysis process should be followed, and in what format will it be delivered ultimately.

There’s another detail: the Gemini Agent can evenwrite its own Skills, saving the methods learned during the work process for use in subsequent tasks.

The fourth is episodic memory.

Used to record which tasks have been completed in the past, along with the corresponding execution experiences.

When the four types of memory are combined, Gemini Agent has the opportunity to gradually get familiar with the user’s working habits, company operations, and team collaboration methods.

For example, when asking it to create a project weekly report for the first time, you may need to explain the report structure, data sources, and workflow.

By the second and third times, it can use the knowledge and methods accumulated before, reducing the amount of information that users need to repeat.

Google also used a metaphor, saying “Gemini will be like a new employee; before starting official work, it will gradually get familiar with you, your tools, and the team.”

Alright.

Now, I can even do the onboarding training by myself.

The only problem is that this employee doesn’t seem to leave voluntarily (here is a smile face.jpg).

In addition, this thing allows you to choose your own large models for work.

Google has clearly stated that the underlying model of Gemini Agent can be selected separately.

At this stage, it can dynamically choose between Google’s own Gemini series and Anthropic’s Claude series models. In the future, more closed-source and open-source models are also planned to be supported.

As for which one to use specifically, Gemini will make a choice between effectiveness, speed, and cost, based on the task requirements.

For simpler tasks, a less costly model can be used; for complex tasks, a more powerful model can be invoked.

In a project, multiple models can be combined to complete different stages.

Google refers to this capability as multi-model orchestration, and it comes with Smart Routing automatic routing functionality.

It is worth noting that Google also emphasized one point—

The underlying model can change, but the context, Skills, and business data accumulated by the Agent can be retained.

It mainly allows it to avoid starting from scratch every time, and can continue from previous tasks, knowledge, and working methods.

This means that even if there are changes in the future model capability rankings, companies can adjust the underlying models without having to completely rebuild all the established workflows.

Google, Meta, OpenAI – a three-way showdown of overseas daily agents?

Google’s entry this time is also a rather subtle timing.

On September 29th, OpenAI just released Dots that can operate 7×24 hours.

Farther ahead, Meta has also launched the personal AI Agent Muse.

Now, Google has joined the battle with Gemini agent, and the three giants have come together in terms of agents.

Let’s take a look at what the other two companies are doing.

Meta’s Muse is more focused on personal life scenarios, and can help users shop, book trips, send emails, and even complete payments.

For example, if you like a product, Muse can help find it, compare options, operate the website, and finally complete the purchase.

Compared to the complex workflows within a company, Muse focuses more on the trivial and time-consuming tasks in an individual's daily life.

This also has something to do with Meta’s existing social products and its large user base.

Letting ordinary users get used to entrusting tasks like shopping and travel to AI is clearly an opportunity that Meta wants to seize.

There are even more similarities between OpenAI’s Dots and Gemini Agent.

Dots has independent cloud computers that operate 7×24 hours, can remember users’ long-term goals and work habits, and can connect to over 4000 applications.

For example, when the developer receives new user feedback, Dots can actively analyze the issue, modify the code, and prepare a PR for review.

For content creators, it can also organize materials and create a draft of the content after new interview records arrive.

However, Dots currently places more emphasis on continuous work around the user’s personal life.

It will gradually understand the user’s preferences, goals, and judgment criteria, and actively look for matters that can be advanced.

OpenAI is also exploring professional Dots tailored for corporate roles, enabling Agents to have independent identities and responsibilities. However, this aspect is currently mainly in the early enterprise pilot stage.

In contrast, what Google is focusing on this time is enabling the Agent to further integrate into an enterprise’s existing organization and office system.

Gemini Agent can utilize the business context in applications such as Gmail, Docs, Sheets, and Calendar to understand what employees are doing.

Companies can also create a Coworker agent with independent email, calendar, accounts, and permissions, allowing it to participate in team collaboration under its own identity.

At the same time, Google has also integrated model selection, enterprise data connectivity, security auditing, and cost control into this system.

Even the competitor Claude can become the underlying model called when the Gemini agent completes its tasks.

From this perspective, the focus of the three companies is quite clear.

One More Thing

By the way, Google has also published other related content in its long article. Those who are interested can check it out by themselves. Cheers!

Gemini has fully joined Google Workspace, enabling work across different applications.

Open Skills, Tools and Enterprise Connector

Introduce Agent for data analysis, finance, and law specialties

Enhance Agent security governance and cost control

【Google Original Text】

https://cloud.google.com/blog/products/ai-machine-learning/welcome-to-gemini-at-work-2026/

Original source

量子位

Content notes

Original publication and rights belong to the source.

Machine translation · Refer to the original