OpenAI Develops Features to Counter Grok Bot, Mulls Response to Meta’s Muse

OpenAI pioneered AI that can take over web browsers and other applications on people’s behalf, and its ChatGPT app brought generative AI to consumers. Now that rivals like SpaceX and Meta are pushing the boundaries by launching personal AI agents that handle multistep tasks, OpenAI is playing catch-up.
For instance, OpenAI is developing features to compete more directly with SpaceX’s recently launched Grok Bot, which helps customers create always-on AI “teammates” to handle routine tasks, manage software, and collaborate in the background, according to a person with knowledge of the effort. OpenAI also has discussed developing a personal AI assistant to compete with one that Meta Platforms recently launched, according to another person with knowledge of the situation.
In developing such products, OpenAI would be repurposing and rebranding some of the existing “agentic” technology it already offers through Codex and ChatGPT. Some customers of these OpenAI apps have already connected them to their email, calendar and other applications so they can automate tasks such as scheduling appointments or communicating with businesses by email, sometimes over multiple days.
And OpenAI has said that since GPT-5 last year, each successive AI model has been designed to tackle longer-horizon tasks, including in those apps.
But many OpenAI users aren’t necessarily aware of these features, and it’s safe to say SpaceX and Meta have stolen some thunder with their respective agent products in recent weeks.
Grok Bot, launched in August, aims to help people quickly set up bots that specialize in different tasks, such as performing market research or developing an ad campaign. Customers can give the bots access to their work applications, such as Gmail and Google Calendar, Notion, Figma, or accounting, finance, AI coding assistants or other software so they can perform tasks within those apps. Grok Bot customers can also tell the AI to watch them perform certain tasks with various apps so it can learn to repeat them on its own.
The bots can report back when they finish a task or need input from the customer. Customers can also require the bots to seek approval before taking actions, such as sending an email, making a purchase or financial transfer, or changing data in another app.
Heavy Advertisement
SpaceX has been positioning Grok Bot—which is included in SuperGrok subscriptions that start at $30 a month—as a kind of AI teammate or digital chief of staff for white-collar workers and small business owners. The company says its bots can operate around the clock, even when the customer’s computer is shut down. It has been advertising Grok Bot heavily across social media, YouTube and podcasts, including the Dwarkesh Podcast focused on AI.
Similarly, Meta this month launched Muse to automate tasks such as communicating with businesses, answering emails, shopping for products and other things people might do from their phone. Meta’s move followed the early rise of personal assistant app Instinct, which is available by invitation only but has more than 100,000 users.
Instinct, which launched before Muse, is trying to raise capital at a $10 billion valuation. It gained attention because it represented a potential breakthrough in AI for consumers after ChatGPT growth slowed dramatically over the past year. (The chatbot earlier this summer surpassed 1 billion weekly active users, a target OpenAI had aimed to hit last year.)
While consumers have gravitated to AI chatbots like ChatGPT and Google’s Gemini to retrieve information and get advice or help with a variety of projects, such chatbots haven’t had as much success in handling more complex, personal tasks such as booking services online.
Tech entrepreneurs and venture capitalists raved online about Instinct’s ability to complete a range of tasks of varying complexity, including managing calendars and ordering supplies for a wedding. Users communicate with Instinct through messaging apps, and can give it access to credit card information and passwords to allow the AI bot to shop or book travel.
It remains to be seen how many consumers are ready to give such AI agents access to personal information so they can perform useful tasks.
Meta has billions of users it can put Muse in front of, and in recent days Meta has promoted the assistant prominently in apps like Messenger. However, Meta’s distribution advantage didn’t work so well in 2024 when the company launched a chatbot that aimed to compete with ChatGPT.
It isn’t clear when OpenAI might launch features to counter Grok Bot and Muse, especially as it balances an endless list of priorities, including ensuring the safety or alignment of its next AI models (and the “swarms” of agents they can power). But it also can’t afford to let competitors grab the AI agent mantle.
Here’s what else is going on…
What OpenAI Researchers See That We Can’t
In the debate over how scared we should be that AI may “kill us all” one day, some have been quick to point out that AI has yet to achieve milestones we might expect before it’s capable of taking over the world: boosting GDP, curing diseases, or automating a substantial number of jobs.
One factor that could help to explain the gap between the public’s experience with AI and the strongly-worded warnings from researchers in the field is the fact that those employees often use new models many months before they reach the masses, as we detailed in this story this morning. The piece focuses on previously unreported talks between Anthropic and OpenAI around testing each others’ models for safety.
Researchers tell us AI has largely automated the process of training new experimental models. Moreover, they also said models within OpenAI today largely write the programs needed to run or train models on graphics processing units (otherwise known as GPU kernels) and the optimizations for those programs.
Engineers can provide the model with a single example of the type of optimization that they want and the AI can run for weeks to implement such optimizations, one OpenAI employee said. This level of AI-powered automation only became possible in the last few months thanks to improvements in models, the employee said.
In another example, the OpenAI employee also said that it’s not uncommon for employees’ agents to work together to solve problems without ever looping in their human users.
Sometimes that can backfire, though: the staffer said that they had noticed times where employees would, for instance, ask their agents to make changes to the company’s codebase and those agents would message other employees on Slack and ask them to fix bugs in the code the agents found, even if their user didn’t ask them to do so or if fixing the bugs weren’t relevant to their work. (Imagine if you asked your agent to complete a task for you and you opened Slack to see your agent chewing out your coworkers for their lazy work!)
A lot of these concerns have to do with reward hacking, or the idea that an AI might find loopholes or take unexpected steps to complete a goal its user gives it. That’s different from an AI acting maliciously on purpose or a bad actor using AI to commit a crime.
OpenAI has taken steps to prevent this type of behavior, we reported. You could imagine, for instance, researchers adding to an AI’s prompt instructions telling it not to extrapolate beyond the exact instructions its user gives it. That’s a careful balancing act, however, because you can also imagine that users might not want to have to write out every single step it wants an agent to take.
There are signs that the rate of AI improvements could accelerate. One reason is OpenAI’s increased access to servers. The ChatGPT-maker’s research team has developed numerous techniques or tricks they have wanted to implement in new models but didn't previously have enough servers to test those theories quickly. Now they do, and what might have previously taken years to pull off now happens in just a week, employees say.
Some of these ideas for improving models are more than a year old, one of the people said. That includes recurrent depth or looping techniques in training and running advanced models.—Stephanie Palazzolo and Amir Efrati
Big Number
OpenAI has told some investors that it expects to burn $278 billion by the end of 2030, the Financial Times reported, as it spends more on cloud computing and chips to run and train its AI. The forecast, based on a recent presentation, is higher than the one it gave investors early this year, when it predicted $180 billion in negative cash flow over the same period.
Overheard
Google acknowledged that its Gemini AI model unexpectedly breached the networks of three outside companies during safety evaluations conducted by third-party testing firm Irregular last May, The Wall Street Journal reported. During the exercise, Gemini gained entry to the companies’ systems by either guessing passwords or finding credentials in public repositories.
A team of cybersecurity researchers with startup Hacktron AI used Anthropic’s Claude to hack into OpenAI, gaining access to a key repository of OpenAI’s software, the researchers revealed Friday. The hack, which occurred on July 25, further demonstrates AI companies’ vulnerability to AI-powered cyberattacks, following the Hugging Face incident, in which OpenAI’s AI models hacked the open source startup as well as OpenAI itself.
The U.S. military nearly intercepted a Chinese vessel after an intelligence report created with AI falsely identified the ship as carrying components of a nuclear program, CNN reported.
Policy Watch
President Donald Trump on Saturday said he was creating a new “AI Force,” that would be similar to his Space Force—the arm of the Defense Department focused on space combat started by Trump during his first administration—and he planned to appoint a new AI czar to lead it. On his Truth Social app, Trump repeated his comments that concerns about AI safety were similar to other topics he considers “hoaxes,” such as global warming, and he again defended data centers.
U.S. Treasury Secretary Scott Bessent said Sunday that the U.S. and China had discussed setting up a mechanism called the U.S.-China AI dialogue to address potential threats. Speaking to reporters at the JPMorgan headquarters in Manhattan at the conclusion of a day of talks with Chinese officials, Bessent said the countries planned to meet again to scope out the notification system, which would create alerts for AI incidents that rise to the national security level.
Ro Khanna, a Democratic congressman for Silicon Valley, has sent letters to Chinese AI companies including Alibaba Group, DeepSeek and Moonshot, requesting their commitments to join American AI labs in a binding international agreement to pace the development of the technology. Khanna, a ranking member of the House Select Committee on China, is also convening an emergency hearing to call for a U.S.-China agreement on AI pacing.
People on the Move
OpenAI has hired Brian McCarthy from SpaceX’s Cursor unit as its new global sales chief, reporting to recently hired Chief Revenue Officer Dali Rajic, the AI firm said Thursday. In addition, two veteran sales executives from Snowflake, including a senior vice president who helped negotiate the firm’s commercial relationship with OpenAI, are leaving to join OpenAI, Snowflake told some employees Thursday.
Alibaba Group has appointed Dayiheng Liu, one of its senior AI researchers, as the head of its Qwen large language models, according to two employees with knowledge of the matter.
Deals and Debuts
See The Information’s Generative AI Database for an exclusive list of private companies and their investors.
Anthropic plans to hold its initial public offering in November rather than October, as had been expected, the Wall Street Journal reported. Advisers to the company say the delay lets it show third-quarter results before listing, after rival OpenAI released its newest model in September.
Nscale, a London company that rents out Nvidia chips and data center capacity to companies training and running AI models, filed to go public on the New York Stock Exchange under the ticker symbol NSCL. Its prospectus shows $140.6 million in revenue for the first six months of 2026, up 1,252% from $10.4 million a year earlier, against a net loss of $1.02 billion, up from $368.9 million.
Crusoe, the data center company behind OpenAI's Abilene, Texas data center, raised $3.9 billion at a $30.9 billion valuation in a Series F funding round led by Atreides Management, Mubadala Capital and Valor Equity Partners, with participation from Founders Fund, GIC, Nvidia, the Qatar Investment Authority, Radical Ventures and TPG.
Manus, an AI agent startup, is in talks with investors including IDG Capital and Boyu Capital to raise $500 million in funding at around a $4 billion valuation, The Wall Street Journal reported.
CoreWeave said on Friday it had priced $3.7 billion in convertible bonds, more than the AI cloud provider had previously said it was targeting.
Angle Health, a company whose AI software lets small businesses build and run their own health insurance plans and helps their employees find care, raised $600 million at a $2.7 billion valuation in a funding round led by Vitruvian Partners.
Naive AI, a Beijing company founded in February by Tsinghua University professor Jifeng Dai that is building a large language model people will be able to download and run themselves, has raised $400 million across three funding rounds from investors including Tencent, IDG Capital, MPCi and HSG, the firm formerly known as Sequoia Capital China.
Vantora, a Santa Monica company formerly called UP.Labs that builds startups to order for large industrial customers, raised more than $100 million in a growth investment from Silversmith Capital Partners.
Mind, a Seattle company whose AI software watches for sensitive corporate data leaking out through AI chatbots, AI agents, email and employee laptops, raised $72 million in a Series B funding round led by Crosspoint Capital Partners.
Raindrop, a San Francisco company whose software watches AI agents while they are working and flags it when they make things up, misuse a tool or start behaving differently after a model upgrade, raised $35 million in a Series A funding round led by CRV.
Unit1, a British company building lifelike digital stand-ins for musicians, raised nearly £15 million, or about $20 million, from investors including Balderton Capital, Mercuri and Paul McGuinness, the music executive who managed U2.
Magentic, a company whose AI software works like a virtual employee to help manufacturers handle purchasing and supply-chain tasks, raised $18 million in a Series A funding round led by Felicis.
Infillion, an advertising technology company, is acquiring Foursquare, the location data company that began in 2009 as an app for checking in to bars and restaurants and now sells data used to target and measure ads. Foursquare will keep operating as its own brand inside Infillion and will keep selling data to customers who use competing ad platforms. Financial terms were not disclosed.
Anthropic has a wet biology lab in the Bay Area where it can use its AI models to run physical experiments, the company has said.
Google Labs announced CC, an experimental agent built for families to help them run their homes.
OpenAI and law firm Cooley co-launched a product, called GO Public, that can draft S-1 filings, the document companies must submit to the Securities and Exchange Commission before they can go public, Cooley said Thursday.
Thank you for reading the AI Agenda Newsletter! I’d love your feedback, ideas and tips: [email protected].
If you think someone else might enjoy this newsletter, please pass it forward or they can sign up here.