How Six AWS Engineers Rebuilt Bedrock to Challenge Microsoft

In early 2025, things were tense inside Amazon’s cloud unit. Bedrock, a new service designed to run AI models, was proving buggy. Customers were getting error messages repeatedly and sometimes waiting weeks for the computing capacity they needed, according to people with knowledge of the customer issues.
As these problems surged, a senior engineering leader, Anthony Liguori, led a team that came up with a plan for a revamped version of Bedrock—known as Project Mantle—designed to fix the throttling and error messages and enable the service to scale to serve larger numbers of users, said two of the people with direct knowledge of the situation. Liguori worked with five other senior engineers to build Mantle using AWS’ in-house AI coding tool, Kiro.
Mantle launched in December last year. A few months later, AWS added OpenAI models to Bedrock’s array of models—which include Anthropic’s Claude, among others. Business is booming.
AWS is now drawing some spending away from Microsoft’s Azure, according to Randall Hunt, chief technology officer at consulting firm Caylent, which advises businesses on how to use AWS. While AWS is bigger than Azure by cloud market share and revenue, Microsoft’s cloud unit previously was able to exclusively offer OpenAI models, giving it an edge over other cloud firms.
Customers are “able to get consistent, reliable performance from Bedrock as opposed to Azure where the performance was a little bit of guesswork,” he said. Hunt said spending on AWS by his financial services clients has grown by five times since May. He says AWS’ AI service offers better performance than rivals like Microsoft.
Amazon said in July that it added more customers to Bedrock in the first half of 2026 than in the service’s first two years, and customers spent more on Bedrock in the second quarter than in all prior quarters combined. Amazon doesn’t break out Bedrock’s revenue, but it reported that AWS’s growth rate accelerated 9 percentage points, to 37%, in the second quarter. Microsoft’s Azure grew 42%, 2 percentage points faster than in the previous quarter.
Microsoft said in a statement that “as AI adoption grows and workloads evolve from model development toward large-scale inference and agent-based applications, Microsoft has continually advanced its AI stack to support changing customer needs and increasing AI demand.”
Bedrock’s early struggles show how the mammoth compute needs of AI have bedeviled even AWS, the industry’s top cloud provider by market share.
An AWS spokesperson said in a statement: “To keep pace with surging demand on Bedrock, we rebuilt its inference engine from the ground up with performance and security in mind. Our approach is clearly resonating with customers. Bedrock is the fastest growing multi-billion-dollar service in AWS history and is used by hundreds of thousands of customers.”
Technical Challenges
AWS launched Bedrock in the fall of 2023, nearly a year after ChatGPT’s introduction in late 2022 kicked off the AI race. Amazon had previously offered AI-powered tools such as SageMaker. But ChatGPT’s popularity caught the company off guard. AWS engineers worked 60-hour weeks to build a service that allowed customers to connect their cloud applications to large language models, according to someone who worked on it.
AWS needed a new service partly because customers couldn’t run leading AI models on its existing systems. Unlike other kinds of software, the models work with proprietary weights that AI firms closely guard and wouldn’t allow customers to access.
But the quick engineering to set up Bedrock led to a product that at times felt “glued together,” according to two people who worked on it later on. Its design meant it didn’t make full use of the AI chip capacity it had, which led to customers getting error messages and hampered AWS’ efforts to scale Bedrock for larger numbers of users, according to one of the people and someone else who worked on it.
An AWS spokesperson said, “We’re constantly making changes to our services and striving for perfection, but [we] also know that operating at our scale will occasionally surface unforeseen challenges, which we identify, communicate to customers and resolve as quickly as possible.”
Meanwhile, capacity was also increasingly a problem for customers. Each user got a base amount of capacity, and if they started to use more, they would be limited until AWS could unlock more compute. For some customers, that wait time could be up to three weeks, a former employee said.
A constant stream of customers throughout 2025 complained about Amazon throttling their access to Anthropic models, according to a person who worked on the product. AWS leaders at one point were concerned that AI coding startup Lovable would defect from Bedrock as a result of its struggles to use Anthropic on Bedrock, The Information previously reported.
In early 2025, an internal report circulated within the AWS sales team said that Bedrock’s technology challenges were undermining efforts to woo customers, according to someone with direct knowledge of the note. Around the time, senior leaders called Bedrock’s capacity problems a “disaster,” The Information reported last year.
Hunt said his clients found that Anthropic ran as much as 63% slower on Bedrock, as measured by tokens per second, than if they went directly to Anthropic in 2024. When businesses go directly to an AI firm, they connect to the model via an application programming interface, although cloud servers host the model. Anthropic used Google Cloud in addition to AWS to offer direct connections via API for its models.
Bedrock’s issues even raised the ire of senior executives in other divisions of Amazon. Dave Treadwell, then a senior vice president in Amazon’s e-commerce division, told AWS leaders in early 2025 that consistent error messages and capacity issues on Bedrock made him want to switch to OpenAI models running directly from the ChatGPT firm, according to two people who heard the comments. OpenAI models that run directly from OpenAI have done so mostly via Microsoft’s Azure cloud.
A person close to Amazon disputes Treadwell, who is now at AWS, made the comment.
Mantle’s Fix
Then came Liguori’s suggestion to create a new version of Bedrock. At AWS’ annual Re:Invent customer conference last December, Dave Brown, then head of its cloud server business, introduced Mantle—a new layer of technology to run Bedrock—and explained that it was designed to let AWS scale the Bedrock service to meet anticipated future demand.
“As models evolve and usage grew, it became clear that inference needed a different architecture,” Brown said in a keynote speech at the event. (Brown recently left AWS and joined Meta Platforms).
Brown also pointed out a shortcoming in Bedrock’s initial design: It was set up using principles applicable for handling requests for standard web services.
But running AI models to power applications—a process known as inference—is more complicated than running older cloud services, where customers store data and pull it out for analysis or other applications. Inference requires finding enough servers to handle workloads even as demand fluctuates constantly. That became especially evident in 2025, when customers were using AI for more complicated tasks.
“Inference works differently than traditional compute patterns we’ve been optimizing for for the past two decades,” Brown said at Re:Invent.
Mantle was designed to accommodate the fact that people using complex AI had “spiky” workloads—ones that take up a lot of compute suddenly and then use very little—at once, whereas older AI only allowed simpler tasks that more standard architecture could handle.
The longer-duration tasks of new AI agents also required a different system to allocate capacity, a spokesperson said.
To solve the problem with spiking at scale, Mantle introduced a new feature where customers could assign different levels of priority to different tasks on Bedrock, so AWS could dedicate more capacity to the ones they cared more about. Mantle also worked to separate the impact of customers’ workloads, so if one customer had a particularly intense workload, it should only affect their capacity and not another customer’s, a system akin to having multiple points of entry at a concert venue, a spokesperson said.
AWS also made Mantle more compatible with an existing AWS service called Journal, which saves an agent’s progress along the way, so customers don’t have to start from scratch if something goes wrong halfway through a long process. This helped Mantle accommodate the longer tasks people were now doing with AI.
“With Mantle, our realization is at the heart of inference is not quite a web service but more of a scheduling system,” said Joe Magerramov, a former AWS vice president who helped create Mantle, on an AWS podcast in January. (Magerramov followed Brown to Meta last month).
Directly Compatible
Mantle was also built to be directly compatible with APIs from OpenAI and Anthropic, a departure from the first version of Bedrock, which required customers to write custom code to use the AI labs’ APIs.
By early 2026, Mantle was up and running. One former employee said the time for customers to get more capacity dropped to two days in late 2025, from three weeks earlier that year. In his investor letter in 2026, Amazon CEO Andy Jassy called out Mantle as the “backbone of our very successful Bedrock service.”
Also later in 2025, Bedrock started running mostly on Amazon’s own AI-specialized Trainium chips, the company said. That gave AWS a new source of computing power as the supply of Nvidia chips continued to face constraints.
To be sure, customers still sometimes encounter delays when using the service. A manager at one AWS customer, which uses Bedrock to access Anthropic Claude models via Mantle, said his firm has experienced intermittent service disruptions in the past few months, the longest of which lasted for around 11 hours.
Half a dozen clients of AI consultancy Loka have switched to running OpenAI models on Bedrock, rather than directly from OpenAI or on Azure, said machine-learning lead Bojan Jakimovski.
There are “some benefits of migrating, [like] fast inference, net more throughput,” said Jakimovski. “The cost is the same—everything is the same.”