Build a Messages API harness that reproduces published DeepSearchQA and BrowseComp scores, using programmatic tool calling, server-side compaction, and task budgets.

Detect safety classifier blocks on Fable 5 and fall back to Opus 4.8 with server-side or SDK-based client-side fallback, including streaming behavior and the new billing changes.

Alexander Bricken
Mahesh Murag

Two async multi-agent patterns — a fixed N-agent team with peer messaging through a shared hub, and dynamically spawned async subagents — reduced to their bare messaging and lifecycle mechanics.

Deploy the research agent from notebook 00 through three tiers of operational maturity (Docker, Modal, Kubernetes) with the same container image and HTTP interface at every tier.

Kevin Tang

Heterogeneous team via the multiagent coordinator config — a coordinator runs three specialists (web-search researcher, file-reading librarian, rules-based pricer) with scoped toolsets to assemble a sales proposal. Covers the multiagent field, the thread_created / thread_message_received event types, and per-role tool scoping.

Build a grade-and-revise loop with Outcomes: a writer drafts a cited research brief, a stateless grader fetches every URL and checks every quote against a rubric, and feedback drives revisions until the brief passes. Covers user.define_outcome, the span.outcome_evaluation_* events, and how to write a rubric the grader can act on.

Give your Claude Managed Agents a Memory store so they learn and remember your users' preferences across multiple interactions.

Build a vulnerability-discovery agent with the Claude Agent SDK that threat-models a C target, hunts memory-safety bugs with built-in file tools, and triages findings into a structured report.

Wire Claude into your on-call flow: when an alert fires, the agent reads logs and runbooks, pinpoints the root cause, opens a fix PR, and waits for your approval before merging.

Build an analyst that turns a CSV into a narrative HTML report with interactive charts, using a sandboxed environment and file mounting.

Mention the bot with a CSV to get an analysis report in-thread, with multi-turn follow-ups on the same session.

Entry-point tutorial for the Claude Managed Agents API. Walks through agent / environment / session creation, file mounts, and the streaming event loop by getting an agent to fix three planted bugs in a calc.py package.

End-to-end production story for Managed Agents — vault-backed MCP credentials, the session.status_idled webhook pattern for human-in-the-loop without long-lived connections, and the resource lifecycle CRUD verbs.

Server-side prompt versioning — create v1, evaluate against a labelled test set, ship v2, detect a regression, roll back by pinning sessions to version 1. Covers agents.update, version pinning on sessions.create, and where the review gate moves when prompts are not code.

Build an agent that autonomously investigates IOCs by querying multiple threat intel sources, cross-referencing findings, mapping to MITRE ATT&CK, and producing structured reports for SIEM and SOAR integration.

List, read, rename, tag, and fork Agent SDK sessions on disk to build a conversation history sidebar without writing a transcript parser.

Build knowledge graphs from unstructured text using Claude for entity extraction, relation mining, deduplication, and multi-hop graph querying.

Compare context engineering strategies for long-running agents and learn when each applies, what it costs, and how they compose.

Port an OpenAI Agents SDK app to the Claude Agent SDK, mapping each primitive (tools, guardrails, sessions, handoffs) through a single expense-approval agent example.

Build an incident response agent with read-write MCP tools for autonomous diagnosis, remediation, and post-mortem documentation.

Manage long-running Claude conversations with instant session memory compaction using background threading and prompt caching.

Reduce latency and token consumption by letting Claude write code that calls tools programmatically in the code execution environment.

Scale Claude applications to thousands of tools using semantic embeddings for dynamic tool discovery.

Manage context limits in long-running agentic workflows by automatically compressing conversation history.

Build a low-latency voice assistant using ElevenLabs for speech-to-text and text-to-speech combined with Claude.

Give Claude a crop tool to zoom into image regions for detailed analysis of charts, documents, and diagrams.

Guide to prompting Claude for distinctive, polished frontend designs avoiding generic aesthetics.

Build financial dashboards and portfolio analytics using Claude's Excel, PowerPoint, PDF skills.

Create, deploy, and manage custom skills extending Claude with specialized organizational workflows.

Create documents, analyze data, automate workflows with Claude's Excel, PowerPoint, PDF skills.

Build a research agent using Claude Code SDK with WebSearch for autonomous research.

Build multi-agent systems with subagents, hooks, output styles, and plan mode features.

Connect agents to external systems via MCP servers for GitHub monitoring and CI workflows.

Run parallel agent evaluations on tools independently from evaluation task files.

Programmatically access and analyze your Claude API usage and cost data via Admin API.

Build AI agents with persistent memory using Claude's memory tool and context editing.

Reduce time-to-first-token by warming cache speculatively while users formulate their queries.

Enable parallel tool calls on Claude 3.7 Sonnet using batch tool meta-pattern workaround.

Use Claude's extended thinking for transparent step-by-step reasoning with budget management.

Combine extended thinking with tools for transparent reasoning during multi-step workflows.

Three simple multi-LLM workflow patterns trading cost or latency for improved performance.

Workflow pattern using one LLM for generation and another for evaluation feedback loop.

Central LLM dynamically delegates tasks to worker LLMs and synthesizes their combined results.

Process large volumes of Claude requests asynchronously with 50% cost reduction using batches.

Convert natural language queries to SQL using RAG, chain-of-thought, and self-improvement techniques.

Improve RAG accuracy by adding context to chunks before embedding with prompt caching.

Step-by-step guide to finetuning Claude 3 Haiku on Amazon Bedrock for custom tasks.

Generate synthetic test cases to evaluate and improve your Claude prompt templates effectively.

Cache and reuse prompt context for cost savings and faster responses with detailed instructions.

Comprehensive guide to summarizing legal documents with evaluation and advanced techniques.

Build and optimize RAG systems with Claude using summary indexing and reranking techniques.

Build classification systems with Claude using RAG and chain-of-thought for insurance tickets.

Control how Claude selects tools using tool_choice parameter for forced or auto selection.

Combine Claude's vision with tools to extract structured data from images like nutrition labels.

Generate longer responses beyond max_tokens limit using prefill technique with message continuation.

Tips and techniques for optimal image processing performance with Claude's vision capabilities.

Create validated tools using Pydantic models for type-safe Claude tool use interactions.

Transcribe audio with Deepgram and generate interview questions using Claude for preparation.

Integrate Wolfram Alpha LLM API as Claude tool for computational queries and answers.

Provide Claude with calculator tool for arithmetic operations and mathematical problem solving.