AWS Machine Learning2d agoenergy 51

How Postman runs Agent Mode for 40 million developers on Amazon Bedrock

image: AWS Machine Learning

Building an AI agent for a demo and operating one for 40 million developers are different engineering problems. Postman set out to build Agent Mode , an AI-native way to work across API testing , documentation , discovery, and implementation. The team expected model quality and prompt design to be the hardest problems. The deeper challenges came from integrating an agent into a mature product with years of interface-driven assumptions, a wide surface area, and specialized concepts. In this post, Postman and AWS describe the architectural patterns that emerged while making a mature product legible to an AI agent. These patterns include controlling tool sprawl, exposing schema-based reads, and treating context rather than capability as the primary bottleneck. We also explain how Agent Mode uses Amazon Bedrock for model flexibility, geographically scoped cross-Region inference, model-dependent zero data retention, and multi-tier prompt caching. Together, these lessons can help teams move production agents beyond prototypes. Why Postman built Agent Mode Agent Mode is Postman’s portal for working with the product in an AI-native way across testing, documentation, discovery, and implementation. Postman has evolved over 11 years, and developers and users learned to locate information through the interface by expanding sidebars, checking tabs, and opening requests. Re-engineering that awareness for an agent surfaced structural assumptions in the product’s APIs, user experience, and distribution of product knowledge. An agent reasons over data rather than navigating a screen. Figure 1 illustrates how Agent Mode works directly against the application. Figure 1: Agent Mode works directly against the Postman application. In this example, it opens a pull request and proposes next steps without requiring the user to navigate through the interface Agent Mode runs on Amazon Bedrock , which provides managed access to foundation models behind the agent. Supporting Postman’s global developer community creates variable, latency-sensitive demand with sharp traffic bursts. With Amazon Bedrock, Postman can scale this production workload without operating its own model-serving infrastructure while retaining flexibility in model selection and control over throughput, geographic processing, and cost. Figure 2 provides a high-level view of the production architecture before the following sections examine its components. Figure 2: Postman Agent Mode combines client-side tools, agent orchestration, purpose-built context, and Amazon Bedrock model inference. Tools are scoped for each task, and user approval remains part of actions that modify application state Human oversight is part of the production design. Agent Mode requires user approval before actions that modify application state. Postman also scopes available tools to the task, selects purpose-built context, and applies model-dependent data-retention settings. These controls reduce unintended actions and unnecessary data exposure, while production testing and monitoring remain necessary. As a responsible AI control, Postman uses Amazon Bedrock Guardrails to redact personally identifiable information before it reaches the underlying large language model (LLM). Enterprise admins can turn this on in Agent Mode’s guardrail settings. Handling tool sprawl In Agent Mode, tools define how the agent acts inside Postman. Early on, the team leaned toward highly atomic tools: small, precise actions such as opening a request, updating one field, or fetching a specific piece of metadata. That approach supported correctness and control in early iterations, but it also revealed several problems. Many real-world workflows require long sequences of tool calls. Even when each step was fast, the overall experience felt slow, because every action had to return to the model before the next one could begin. Users watched the agent step through actions they had mentally grouped as a single operation. In Postman’s testing, tool-selection errors increased once the visible toolset exceeded approximately 40 tools. The agent could call nonexistent tools, pass incorrect arguments despite valid schemas, or select tools that seemed semantically reasonable but were wrong in context. Larger or newer models reduced this behavior but didn’t remove it. Past a certain toolset size, exposing more tools can reduce agent effectiveness. The current architecture selects tools based on need and context and isolates individual execution threads. The model sees only the tools relevant to the current task. Figure 3 illustrates this dynamic selection process. Figure 3: The root agent queries a vector database of tool embeddings and narrows more than 170 tools to approximately 15 relevant to the request. It then hands those tools to a context-isolated sub-agent, so the model sees only the tools needed for the task A subtler problem was that many client APIs were implicitly coupled to interface state. Tools that modified requests needed certain elements to be open, while other tools opened new tabs as side effects. The agent had to open a request tab to read it, mimicking interface interactions instead of reasoning about data. Postman is actively decoupling tools from tabs, and its Native Git feature makes extensive use of this approach. For example, Agent Mode can now send requests in the background without an open tab, although user approval is still required. Builder takeaway: Treat your tool catalog as part of the context budget. Dynamically scope the tools exposed to the model per task, and decouple “what the agent can do” from “what the UI happens to have open.” Exposing schema-based reads For products such as the API Catalog, Postman consolidated multiple narrow views into a single query tool. These products expose structured data such as service uptime, test results, and endpoint response times across many services. Given the schemas of the underlying ClickHouse tables, the agent can generate complex queries with joins and WHERE clauses. This substantially reduces the number of distinct tools needed to answer an analysis question: SELECT toString(service_id) AS service_id, countMerge(total_events_state) AS total_requests, countMerge(error_events_state) AS total_errors, round(countMerge(error_events_state) * 100.0 / countMerge(total_events_state), 4) AS error_rate_pct, avgMerge(avg_latency_state) AS avg_latency_ms, quantileMerge(0.95)(p95_latency_state) AS p95_latency_ms FROM http_events_summary_1d WHERE service_id IN ('...list of service IDs') AND bucket_1d >= today() - 7 GROUP BY service_id HAVING p95_latency_ms < 100 AND total_requests > 0 ORDER BY error_rate_pct DESC; With this approach, the engineering job shifts from building a tool per question to modeling the data well once . The agent can then generate a far wider variety of queries than the team could ever have enumerated as individual tools. Builder takeaway: Where you have well-structured data, give the agent schema-aware read access to a query engine instead of a proliferation of single-purpose read tools. You trade tool count for data modeling, which produces a better scaling curve. Context was the real bottleneck Postman initially assumed missing tools would be the biggest blocker. In practice, missing or incomplete context caused more failures than missing capabilities. Context is the agent’s understanding of where the user is in Postman, which entities are active, and what state has already been established. When that context was wrong or absent, even correct tools became ineffective. Figure 4 distinguishes the two forms of context supplied to the agent. Figure 4: Two kinds of context feed the agent. Broad, shallow background context is gathered automatically and minified for the prompt. Deep, focused selected context is chosen by the user and routed through a dedicated handler for each entity type. Each handler distills the entity into the information the agent needs The challenge was structural

read the original at AWS Machine Learning →