Research agent stack
An agent that searches the web and the scholarly record, reads full sources, runs analysis code in a sandbox and cites what it found, with tracing so you can audit how it got there.
Use case
Answer open research questions with cited sources by searching the web and academic indexes, reading full texts, running analysis in a sandbox and logging every step for review.
Components
| Role | Component | Why |
|---|---|---|
| Agent harness | Claude Agent SDK | Gives you Claude Code's agent loop, tools and context management as a Python or TypeScript library, with MCP support, so the research loop runs in your own app or CI. |
| Web search | Exa API | Search built for agents on Exa's own index, with page contents in the same call. A hosted MCP server at mcp.exa.ai/mcp works without an API key to start. |
| Web search (alternative) | Tavily API | Search plus extract, crawl and map endpoints returning LLM-ready content; a good second source when you want to cross-check results across indexes. |
| Page reading | Firecrawl API | Scrapes pages to clean Markdown, with browser actions for dynamic pages, plus crawl and map endpoints; open source (AGPL-3.0) if you need to self-host. |
| Quick page reading | Jina Reader | Prefix any URL with r.jina.ai to get Markdown. Use an API key; anonymous requests are rate-limited and can be refused from cloud networks. |
| Scholarly metadata | OpenAlex API | Open catalog of works, authors, institutions and citations with CC0 data. A free account's API key includes $1 of usage per day; paid plans add more. |
| Papers and citations | Semantic Scholar Academic Graph API | Paper search, citation graphs and author data. Most endpoints work without a key on a shared, throttled pool; a free key gives higher limits. |
| Preprints | arXiv API | Search and metadata for arXiv preprints. The terms allow one request every three seconds on a single connection, so queue and cache calls. |
| Browser | Playwright MCP | A real browser for pages that block simple fetches or need clicking; works from accessibility snapshots rather than screenshots. |
| Analysis sandbox | E2B | Runs the agent's pandas or plotting code in a disposable microVM, with the network turned off once data is loaded. |
| Tracing and evals | Langfuse | Open-source tracing of every model call and tool call (MIT outside its enterprise directories), so you can check which sources an answer really came from and build evals. |
How the pieces fit
A research agent runs the same loop over and over: plan the question, search, read, take notes, check, write. The stack above maps one component to each step.
- Search wide, then deep. Start with a web search API (Exa, with Tavily, Brave Search or Parallel as alternatives) for recent and general sources, and the scholarly APIs (OpenAlex, Semantic Scholar, arXiv) for peer-reviewed work and citation trails. Search results are snippets; don't let the agent cite a snippet.
- Read the source. Fetch full pages as Markdown with Firecrawl or Jina Reader, and fall back to a real browser (Playwright MCP) when a page needs JavaScript or interaction.
- Compute in a sandbox. When the question needs numbers, have the agent write analysis code and run it in E2B, not on your machine. Load the data, then cut the network. The sandboxing guide explains why.
- Trace everything. Send every model call and tool call to Langfuse. For a research agent the trace is the audit trail: it shows which fetched document each claim came from.
The harness here is the Claude Agent SDK; the OpenAI Agents SDK or LangGraph work the same way if you prefer them. If your harness speaks MCP, most of these components can be connected as MCP servers instead of custom tools; see the remote MCP guide.
Practical limits to plan for
- Rate limits. arXiv asks for at most one request every three seconds across all your machines. Semantic Scholar's unauthenticated pool is shared and throttled. OpenAlex's free tier is a daily usage budget tied to an API key. Put a queue and a cache in front of the scholarly APIs, and give each its own key.
- Prompt injection. Every page the agent reads is untrusted input. Keep the browser and fetch tools away from credentials, and don't give the same agent a tool that can send email or spend money.
- Citation hygiene. Require a URL or DOI for every claim, and have a final pass re-fetch each cited source and check that the quoted text is there. Search snippets and model memory are not sources.
- Memory across sessions. For long projects, keep notes in files the agent re-reads, or add a memory layer; see giving an agent long-term memory.
To choose between the search APIs, see the web search APIs comparison.
Directory entries in this stack
- Claude Agent SDK: Anthropic's SDK for building production agents with Claude Code as a library, in Python and TypeScript.
- Exa API: Search and retrieval API built for AI agents, backed by Exa's own continuously updated web index and search models.
- Exa MCP Server: Official Exa MCP server connecting agents to Exa for web search, content fetching and multi-step research.
- Tavily API: Real-time web layer for AI agents: search, extract, crawl, map and cited research returned as clean LLM-ready content.
- Firecrawl API: Web data API for AI agents: search the web, scrape any page to clean data and interact with it through one API.
- Jina Reader: Converts any URL into LLM-friendly Markdown by prefixing it with https://r.jina.ai/.
- OpenAlex API: Index of half a billion scholarly works with their authors, institutions, sources and topics, available via API and bulk download.
- Semantic Scholar Academic Graph API: API over Semantic Scholar's academic graph of papers, authors and citations; an API key unlocks some endpoints and higher rate limits.
- arXiv API: Programmatic search and metadata access to arXiv preprints via an Atom-based query interface.
- Playwright MCP: Microsoft's MCP server for browser automation with Playwright, letting LLMs act on web pages via structured accessibility snapshots instead of screenshots.
- E2B: Open-source, secure cloud sandboxes for AI agents: an isolated machine per agent to run code, browse and use tools.
- Langfuse: Open-source agent evals and observability platform: trace, evaluate and improve LLM applications and agents.
- Brave Search API: Web search API on Brave's independent index of 40+ billion pages, with specialized endpoints for search, AI grounding and more.
- Parallel Web APIs: Web APIs for AI agents (Search, Extract, Task/Deep Research, FindAll, Monitor) for research, extraction and continuous monitoring.
- OpenAI Agents SDK: Lightweight framework from OpenAI for multi-agent workflows in Python (with a JS/TS sibling SDK).
- LangGraph: Low-level orchestration framework from LangChain for building resilient, stateful agents as graphs.
Related
- Web search APIs for agents (Comparison)
- Run agent-generated code safely in a sandbox (Guide)
- Code sandboxes for AI agents (Comparison)
- Give an agent long-term memory (Guide)
- Connect an agent to a remote MCP server (Guide)
Sources
- Claude Code docs, Run Claude Code programmatically (Agent SDK), accessed
- Exa docs, Exa MCP, accessed
- Jina Reader, accessed
- OpenAlex Help, Pricing overview, accessed
- Semantic Scholar API, accessed
- arXiv, Terms of Use for arXiv APIs, accessed
- Langfuse LICENSE, accessed
- E2B docs, Internet access, accessed