{
  "type": "stack",
  "id": "research-agent",
  "title": "Research agent stack",
  "summary": "An agent that searches the web and the scholarly record, reads full sources, runs analysis code in a sandbox and cites what it found, with tracing so you can audit how it got there.",
  "author": "Agentica Author",
  "tags": [
    "research",
    "web-search",
    "academic",
    "agents",
    "citations"
  ],
  "published": "2026-10-02",
  "last_verified": "2026-10-02",
  "entries": [
    "claude-agent-sdk",
    "exa-api",
    "exa-mcp-server",
    "tavily-api",
    "firecrawl-api",
    "jina-reader",
    "openalex-api",
    "semantic-scholar-api",
    "arxiv-api",
    "playwright-mcp",
    "e2b",
    "langfuse",
    "brave-search-api",
    "parallel-api",
    "openai-agents-sdk",
    "langgraph"
  ],
  "links": {
    "html": "https://indexagentica.com/stacks/research-agent/",
    "markdown": "https://indexagentica.com/stacks/research-agent.md",
    "json": "https://indexagentica.com/api/longform/stacks/research-agent.json",
    "source": "https://github.com/Drudley/indexagentica/blob/main/content-long/stacks/research-agent.md"
  },
  "status": "published",
  "stack": {
    "use_case": "Answer open research questions with cited sources by searching the web and academic indexes, reading full texts, running analysis in a sandbox and logging every step for review.",
    "components": [
      {
        "role": "Agent harness",
        "entry": "claude-agent-sdk",
        "name": "Claude Agent SDK",
        "url": "https://indexagentica.com/entries/claude-agent-sdk/",
        "why": "Gives you Claude Code's agent loop, tools and context management as a Python or TypeScript library, with MCP support, so the research loop runs in your own app or CI."
      },
      {
        "role": "Web search",
        "entry": "exa-api",
        "name": "Exa API",
        "url": "https://indexagentica.com/entries/exa-api/",
        "why": "Search built for agents on Exa's own index, with page contents in the same call. A hosted MCP server at mcp.exa.ai/mcp works without an API key to start."
      },
      {
        "role": "Web search (alternative)",
        "entry": "tavily-api",
        "name": "Tavily API",
        "url": "https://indexagentica.com/entries/tavily-api/",
        "why": "Search plus extract, crawl and map endpoints returning LLM-ready content; a good second source when you want to cross-check results across indexes."
      },
      {
        "role": "Page reading",
        "entry": "firecrawl-api",
        "name": "Firecrawl API",
        "url": "https://indexagentica.com/entries/firecrawl-api/",
        "why": "Scrapes pages to clean Markdown, with browser actions for dynamic pages, plus crawl and map endpoints; open source (AGPL-3.0) if you need to self-host."
      },
      {
        "role": "Quick page reading",
        "entry": "jina-reader",
        "name": "Jina Reader",
        "url": "https://indexagentica.com/entries/jina-reader/",
        "why": "Prefix any URL with r.jina.ai to get Markdown. Use an API key; anonymous requests are rate-limited and can be refused from cloud networks."
      },
      {
        "role": "Scholarly metadata",
        "entry": "openalex-api",
        "name": "OpenAlex API",
        "url": "https://indexagentica.com/entries/openalex-api/",
        "why": "Open catalog of works, authors, institutions and citations with CC0 data. A free account's API key includes $1 of usage per day; paid plans add more."
      },
      {
        "role": "Papers and citations",
        "entry": "semantic-scholar-api",
        "name": "Semantic Scholar Academic Graph API",
        "url": "https://indexagentica.com/entries/semantic-scholar-api/",
        "why": "Paper search, citation graphs and author data. Most endpoints work without a key on a shared, throttled pool; a free key gives higher limits."
      },
      {
        "role": "Preprints",
        "entry": "arxiv-api",
        "name": "arXiv API",
        "url": "https://indexagentica.com/entries/arxiv-api/",
        "why": "Search and metadata for arXiv preprints. The terms allow one request every three seconds on a single connection, so queue and cache calls."
      },
      {
        "role": "Browser",
        "entry": "playwright-mcp",
        "name": "Playwright MCP",
        "url": "https://indexagentica.com/entries/playwright-mcp/",
        "why": "A real browser for pages that block simple fetches or need clicking; works from accessibility snapshots rather than screenshots."
      },
      {
        "role": "Analysis sandbox",
        "entry": "e2b",
        "name": "E2B",
        "url": "https://indexagentica.com/entries/e2b/",
        "why": "Runs the agent's pandas or plotting code in a disposable microVM, with the network turned off once data is loaded."
      },
      {
        "role": "Tracing and evals",
        "entry": "langfuse",
        "name": "Langfuse",
        "url": "https://indexagentica.com/entries/langfuse/",
        "why": "Open-source tracing of every model call and tool call (MIT outside its enterprise directories), so you can check which sources an answer really came from and build evals."
      }
    ]
  },
  "entries_detail": [
    {
      "id": "claude-agent-sdk",
      "name": "Claude Agent SDK",
      "summary": "Anthropic's SDK for building production agents with Claude Code as a library, in Python and TypeScript.",
      "url": "https://indexagentica.com/entries/claude-agent-sdk/",
      "json": "https://indexagentica.com/api/entries/claude-agent-sdk.json"
    },
    {
      "id": "exa-api",
      "name": "Exa API",
      "summary": "Search and retrieval API built for AI agents, backed by Exa's own continuously updated web index and search models.",
      "url": "https://indexagentica.com/entries/exa-api/",
      "json": "https://indexagentica.com/api/entries/exa-api.json"
    },
    {
      "id": "exa-mcp-server",
      "name": "Exa MCP Server",
      "summary": "Official Exa MCP server connecting agents to Exa for web search, content fetching and multi-step research.",
      "url": "https://indexagentica.com/entries/exa-mcp-server/",
      "json": "https://indexagentica.com/api/entries/exa-mcp-server.json"
    },
    {
      "id": "tavily-api",
      "name": "Tavily API",
      "summary": "Real-time web layer for AI agents: search, extract, crawl, map and cited research returned as clean LLM-ready content.",
      "url": "https://indexagentica.com/entries/tavily-api/",
      "json": "https://indexagentica.com/api/entries/tavily-api.json"
    },
    {
      "id": "firecrawl-api",
      "name": "Firecrawl API",
      "summary": "Web data API for AI agents: search the web, scrape any page to clean data and interact with it through one API.",
      "url": "https://indexagentica.com/entries/firecrawl-api/",
      "json": "https://indexagentica.com/api/entries/firecrawl-api.json"
    },
    {
      "id": "jina-reader",
      "name": "Jina Reader",
      "summary": "Converts any URL into LLM-friendly Markdown by prefixing it with https://r.jina.ai/.",
      "url": "https://indexagentica.com/entries/jina-reader/",
      "json": "https://indexagentica.com/api/entries/jina-reader.json"
    },
    {
      "id": "openalex-api",
      "name": "OpenAlex API",
      "summary": "Index of half a billion scholarly works with their authors, institutions, sources and topics, available via API and bulk download.",
      "url": "https://indexagentica.com/entries/openalex-api/",
      "json": "https://indexagentica.com/api/entries/openalex-api.json"
    },
    {
      "id": "semantic-scholar-api",
      "name": "Semantic Scholar Academic Graph API",
      "summary": "API over Semantic Scholar's academic graph of papers, authors and citations; an API key unlocks some endpoints and higher rate limits.",
      "url": "https://indexagentica.com/entries/semantic-scholar-api/",
      "json": "https://indexagentica.com/api/entries/semantic-scholar-api.json"
    },
    {
      "id": "arxiv-api",
      "name": "arXiv API",
      "summary": "Programmatic search and metadata access to arXiv preprints via an Atom-based query interface.",
      "url": "https://indexagentica.com/entries/arxiv-api/",
      "json": "https://indexagentica.com/api/entries/arxiv-api.json"
    },
    {
      "id": "playwright-mcp",
      "name": "Playwright MCP",
      "summary": "Microsoft's MCP server for browser automation with Playwright, letting LLMs act on web pages via structured accessibility snapshots instead of screenshots.",
      "url": "https://indexagentica.com/entries/playwright-mcp/",
      "json": "https://indexagentica.com/api/entries/playwright-mcp.json"
    },
    {
      "id": "e2b",
      "name": "E2B",
      "summary": "Open-source, secure cloud sandboxes for AI agents: an isolated machine per agent to run code, browse and use tools.",
      "url": "https://indexagentica.com/entries/e2b/",
      "json": "https://indexagentica.com/api/entries/e2b.json"
    },
    {
      "id": "langfuse",
      "name": "Langfuse",
      "summary": "Open-source agent evals and observability platform: trace, evaluate and improve LLM applications and agents.",
      "url": "https://indexagentica.com/entries/langfuse/",
      "json": "https://indexagentica.com/api/entries/langfuse.json"
    },
    {
      "id": "brave-search-api",
      "name": "Brave Search API",
      "summary": "Web search API on Brave's independent index of 40+ billion pages, with specialized endpoints for search, AI grounding and more.",
      "url": "https://indexagentica.com/entries/brave-search-api/",
      "json": "https://indexagentica.com/api/entries/brave-search-api.json"
    },
    {
      "id": "parallel-api",
      "name": "Parallel Web APIs",
      "summary": "Web APIs for AI agents (Search, Extract, Task/Deep Research, FindAll, Monitor) for research, extraction and continuous monitoring.",
      "url": "https://indexagentica.com/entries/parallel-api/",
      "json": "https://indexagentica.com/api/entries/parallel-api.json"
    },
    {
      "id": "openai-agents-sdk",
      "name": "OpenAI Agents SDK",
      "summary": "Lightweight framework from OpenAI for multi-agent workflows in Python (with a JS/TS sibling SDK).",
      "url": "https://indexagentica.com/entries/openai-agents-sdk/",
      "json": "https://indexagentica.com/api/entries/openai-agents-sdk.json"
    },
    {
      "id": "langgraph",
      "name": "LangGraph",
      "summary": "Low-level orchestration framework from LangChain for building resilient, stateful agents as graphs.",
      "url": "https://indexagentica.com/entries/langgraph/",
      "json": "https://indexagentica.com/api/entries/langgraph.json"
    }
  ],
  "related": [
    {
      "type": "comparison",
      "id": "web-search-apis",
      "title": "Web search APIs for agents",
      "url": "https://indexagentica.com/compare/web-search-apis/",
      "json": "https://indexagentica.com/api/longform/compare/web-search-apis.json"
    },
    {
      "type": "guide",
      "id": "run-untrusted-code-in-a-sandbox",
      "title": "Run agent-generated code safely in a sandbox",
      "url": "https://indexagentica.com/guides/run-untrusted-code-in-a-sandbox/",
      "json": "https://indexagentica.com/api/longform/guides/run-untrusted-code-in-a-sandbox.json"
    },
    {
      "type": "comparison",
      "id": "code-sandboxes",
      "title": "Code sandboxes for AI agents",
      "url": "https://indexagentica.com/compare/code-sandboxes/",
      "json": "https://indexagentica.com/api/longform/compare/code-sandboxes.json"
    },
    {
      "type": "guide",
      "id": "give-an-agent-long-term-memory",
      "title": "Give an agent long-term memory",
      "url": "https://indexagentica.com/guides/give-an-agent-long-term-memory/",
      "json": "https://indexagentica.com/api/longform/guides/give-an-agent-long-term-memory.json"
    },
    {
      "type": "guide",
      "id": "connect-an-agent-to-a-remote-mcp-server",
      "title": "Connect an agent to a remote MCP server",
      "url": "https://indexagentica.com/guides/connect-an-agent-to-a-remote-mcp-server/",
      "json": "https://indexagentica.com/api/longform/guides/connect-an-agent-to-a-remote-mcp-server.json"
    }
  ],
  "sources": [
    {
      "title": "Claude Code docs, Run Claude Code programmatically (Agent SDK)",
      "url": "https://code.claude.com/docs/en/headless",
      "accessed": "2026-10-02"
    },
    {
      "title": "Exa docs, Exa MCP",
      "url": "https://exa.ai/docs/get-started/exa-mcp",
      "accessed": "2026-10-02"
    },
    {
      "title": "Jina Reader",
      "url": "https://jina.ai/reader",
      "accessed": "2026-10-02"
    },
    {
      "title": "OpenAlex Help, Pricing overview",
      "url": "https://help.openalex.org/access/pricing",
      "accessed": "2026-10-02"
    },
    {
      "title": "Semantic Scholar API",
      "url": "https://www.semanticscholar.org/product/api",
      "accessed": "2026-10-02"
    },
    {
      "title": "arXiv, Terms of Use for arXiv APIs",
      "url": "https://info.arxiv.org/help/api/tou.html",
      "accessed": "2026-10-02"
    },
    {
      "title": "Langfuse LICENSE",
      "url": "https://github.com/langfuse/langfuse/blob/main/LICENSE",
      "accessed": "2026-10-02"
    },
    {
      "title": "E2B docs, Internet access",
      "url": "https://docs.e2b.dev/network/internet-access",
      "accessed": "2026-10-02"
    }
  ],
  "front_matter": {
    "id": "research-agent",
    "type": "stack",
    "title": "Research agent stack",
    "summary": "An agent that searches the web and the scholarly record, reads full sources, runs analysis code in a sandbox and cites what it found, with tracing so you can audit how it got there.",
    "description": "A component list for a deep-research style agent, with the reason each piece is there and the limits that matter when you wire it up (rate limits, free tiers, authentication). Every component is swappable; the alternatives named in the notes are directory entries too.",
    "author": "Agentica Author",
    "use_case": "Answer open research questions with cited sources by searching the web and academic indexes, reading full texts, running analysis in a sandbox and logging every step for review.",
    "components": [
      {
        "role": "Agent harness",
        "entry": "claude-agent-sdk",
        "why": "Gives you Claude Code's agent loop, tools and context management as a Python or TypeScript library, with MCP support, so the research loop runs in your own app or CI."
      },
      {
        "role": "Web search",
        "entry": "exa-api",
        "why": "Search built for agents on Exa's own index, with page contents in the same call. A hosted MCP server at mcp.exa.ai/mcp works without an API key to start."
      },
      {
        "role": "Web search (alternative)",
        "entry": "tavily-api",
        "why": "Search plus extract, crawl and map endpoints returning LLM-ready content; a good second source when you want to cross-check results across indexes."
      },
      {
        "role": "Page reading",
        "entry": "firecrawl-api",
        "why": "Scrapes pages to clean Markdown, with browser actions for dynamic pages, plus crawl and map endpoints; open source (AGPL-3.0) if you need to self-host."
      },
      {
        "role": "Quick page reading",
        "entry": "jina-reader",
        "why": "Prefix any URL with r.jina.ai to get Markdown. Use an API key; anonymous requests are rate-limited and can be refused from cloud networks."
      },
      {
        "role": "Scholarly metadata",
        "entry": "openalex-api",
        "why": "Open catalog of works, authors, institutions and citations with CC0 data. A free account's API key includes $1 of usage per day; paid plans add more."
      },
      {
        "role": "Papers and citations",
        "entry": "semantic-scholar-api",
        "why": "Paper search, citation graphs and author data. Most endpoints work without a key on a shared, throttled pool; a free key gives higher limits."
      },
      {
        "role": "Preprints",
        "entry": "arxiv-api",
        "why": "Search and metadata for arXiv preprints. The terms allow one request every three seconds on a single connection, so queue and cache calls."
      },
      {
        "role": "Browser",
        "entry": "playwright-mcp",
        "why": "A real browser for pages that block simple fetches or need clicking; works from accessibility snapshots rather than screenshots."
      },
      {
        "role": "Analysis sandbox",
        "entry": "e2b",
        "why": "Runs the agent's pandas or plotting code in a disposable microVM, with the network turned off once data is loaded."
      },
      {
        "role": "Tracing and evals",
        "entry": "langfuse",
        "why": "Open-source tracing of every model call and tool call (MIT outside its enterprise directories), so you can check which sources an answer really came from and build evals."
      }
    ],
    "tags": [
      "research",
      "web-search",
      "academic",
      "agents",
      "citations"
    ],
    "entries": [
      "claude-agent-sdk",
      "exa-api",
      "exa-mcp-server",
      "tavily-api",
      "firecrawl-api",
      "jina-reader",
      "openalex-api",
      "semantic-scholar-api",
      "arxiv-api",
      "playwright-mcp",
      "e2b",
      "langfuse",
      "brave-search-api",
      "parallel-api",
      "openai-agents-sdk",
      "langgraph"
    ],
    "sources": [
      {
        "title": "Claude Code docs, Run Claude Code programmatically (Agent SDK)",
        "url": "https://code.claude.com/docs/en/headless",
        "accessed": "2026-10-02"
      },
      {
        "title": "Exa docs, Exa MCP",
        "url": "https://exa.ai/docs/get-started/exa-mcp",
        "accessed": "2026-10-02"
      },
      {
        "title": "Jina Reader",
        "url": "https://jina.ai/reader",
        "accessed": "2026-10-02"
      },
      {
        "title": "OpenAlex Help, Pricing overview",
        "url": "https://help.openalex.org/access/pricing",
        "accessed": "2026-10-02"
      },
      {
        "title": "Semantic Scholar API",
        "url": "https://www.semanticscholar.org/product/api",
        "accessed": "2026-10-02"
      },
      {
        "title": "arXiv, Terms of Use for arXiv APIs",
        "url": "https://info.arxiv.org/help/api/tou.html",
        "accessed": "2026-10-02"
      },
      {
        "title": "Langfuse LICENSE",
        "url": "https://github.com/langfuse/langfuse/blob/main/LICENSE",
        "accessed": "2026-10-02"
      },
      {
        "title": "E2B docs, Internet access",
        "url": "https://docs.e2b.dev/network/internet-access",
        "accessed": "2026-10-02"
      }
    ],
    "related": [
      "web-search-apis",
      "run-untrusted-code-in-a-sandbox",
      "code-sandboxes",
      "give-an-agent-long-term-memory",
      "connect-an-agent-to-a-remote-mcp-server"
    ],
    "last_verified": "2026-10-02",
    "published": "2026-10-02"
  },
  "markdown": "\n## How the pieces fit\n\nA research agent runs the same loop over and over: plan the question, search, read, take notes, check, write. The stack above maps one component to each step.\n\n1. **Search wide, then deep.** Start with a web search API ([Exa](https://indexagentica.com/entries/exa-api/), with [Tavily](https://indexagentica.com/entries/tavily-api/), [Brave Search](https://indexagentica.com/entries/brave-search-api/) or [Parallel](https://indexagentica.com/entries/parallel-api/) as alternatives) for recent and general sources, and the scholarly APIs ([OpenAlex](https://indexagentica.com/entries/openalex-api/), [Semantic Scholar](https://indexagentica.com/entries/semantic-scholar-api/), [arXiv](https://indexagentica.com/entries/arxiv-api/)) for peer-reviewed work and citation trails. Search results are snippets; don't let the agent cite a snippet.\n2. **Read the source.** Fetch full pages as Markdown with [Firecrawl](https://indexagentica.com/entries/firecrawl-api/) or [Jina Reader](https://indexagentica.com/entries/jina-reader/), and fall back to a real browser ([Playwright MCP](https://indexagentica.com/entries/playwright-mcp/)) when a page needs JavaScript or interaction.\n3. **Compute in a sandbox.** When the question needs numbers, have the agent write analysis code and run it in [E2B](https://indexagentica.com/entries/e2b/), not on your machine. Load the data, then cut the network. The [sandboxing guide](https://indexagentica.com/guides/run-untrusted-code-in-a-sandbox/) explains why.\n4. **Trace everything.** Send every model call and tool call to [Langfuse](https://indexagentica.com/entries/langfuse/). For a research agent the trace is the audit trail: it shows which fetched document each claim came from.\n\nThe harness here is the [Claude Agent SDK](https://indexagentica.com/entries/claude-agent-sdk/); the [OpenAI Agents SDK](https://indexagentica.com/entries/openai-agents-sdk/) or [LangGraph](https://indexagentica.com/entries/langgraph/) work the same way if you prefer them. If your harness speaks MCP, most of these components can be connected as MCP servers instead of custom tools; see the [remote MCP guide](https://indexagentica.com/guides/connect-an-agent-to-a-remote-mcp-server/).\n\n## Practical limits to plan for\n\n- **Rate limits.** arXiv asks for at most one request every three seconds across all your machines. Semantic Scholar's unauthenticated pool is shared and throttled. OpenAlex's free tier is a daily usage budget tied to an API key. Put a queue and a cache in front of the scholarly APIs, and give each its own key.\n- **Prompt injection.** Every page the agent reads is untrusted input. Keep the browser and fetch tools away from credentials, and don't give the same agent a tool that can send email or spend money.\n- **Citation hygiene.** Require a URL or DOI for every claim, and have a final pass re-fetch each cited source and check that the quoted text is there. Search snippets and model memory are not sources.\n- **Memory across sessions.** For long projects, keep notes in files the agent re-reads, or add a memory layer; see [giving an agent long-term memory](https://indexagentica.com/guides/give-an-agent-long-term-memory/).\n\nTo choose between the search APIs, see the [web search APIs comparison](https://indexagentica.com/compare/web-search-apis/).\n",
  "raw": "---\nid: research-agent\ntype: stack\ntitle: Research agent stack\nsummary: An agent that searches the web and the scholarly record, reads full sources, runs analysis code in a sandbox and cites what it found, with tracing so you can audit how it got there.\ndescription: \"A component list for a deep-research style agent, with the reason each piece is there and the limits that matter when you wire it up (rate limits, free tiers, authentication). Every component is swappable; the alternatives named in the notes are directory entries too.\"\nauthor: Agentica Author\nuse_case: Answer open research questions with cited sources by searching the web and academic indexes, reading full texts, running analysis in a sandbox and logging every step for review.\ncomponents:\n  - role: Agent harness\n    entry: claude-agent-sdk\n    why: \"Gives you Claude Code's agent loop, tools and context management as a Python or TypeScript library, with MCP support, so the research loop runs in your own app or CI.\"\n  - role: Web search\n    entry: exa-api\n    why: \"Search built for agents on Exa's own index, with page contents in the same call. A hosted MCP server at mcp.exa.ai/mcp works without an API key to start.\"\n  - role: Web search (alternative)\n    entry: tavily-api\n    why: \"Search plus extract, crawl and map endpoints returning LLM-ready content; a good second source when you want to cross-check results across indexes.\"\n  - role: Page reading\n    entry: firecrawl-api\n    why: \"Scrapes pages to clean Markdown, with browser actions for dynamic pages, plus crawl and map endpoints; open source (AGPL-3.0) if you need to self-host.\"\n  - role: Quick page reading\n    entry: jina-reader\n    why: \"Prefix any URL with r.jina.ai to get Markdown. Use an API key; anonymous requests are rate-limited and can be refused from cloud networks.\"\n  - role: Scholarly metadata\n    entry: openalex-api\n    why: \"Open catalog of works, authors, institutions and citations with CC0 data. A free account's API key includes $1 of usage per day; paid plans add more.\"\n  - role: Papers and citations\n    entry: semantic-scholar-api\n    why: \"Paper search, citation graphs and author data. Most endpoints work without a key on a shared, throttled pool; a free key gives higher limits.\"\n  - role: Preprints\n    entry: arxiv-api\n    why: \"Search and metadata for arXiv preprints. The terms allow one request every three seconds on a single connection, so queue and cache calls.\"\n  - role: Browser\n    entry: playwright-mcp\n    why: \"A real browser for pages that block simple fetches or need clicking; works from accessibility snapshots rather than screenshots.\"\n  - role: Analysis sandbox\n    entry: e2b\n    why: \"Runs the agent's pandas or plotting code in a disposable microVM, with the network turned off once data is loaded.\"\n  - role: Tracing and evals\n    entry: langfuse\n    why: \"Open-source tracing of every model call and tool call (MIT outside its enterprise directories), so you can check which sources an answer really came from and build evals.\"\ntags: [research, web-search, academic, agents, citations]\nentries: [claude-agent-sdk, exa-api, exa-mcp-server, tavily-api, firecrawl-api, jina-reader, openalex-api, semantic-scholar-api, arxiv-api, playwright-mcp, e2b, langfuse, brave-search-api, parallel-api, openai-agents-sdk, langgraph]\nsources:\n  - title: Claude Code docs, Run Claude Code programmatically (Agent SDK)\n    url: https://code.claude.com/docs/en/headless\n    accessed: 2026-10-02\n  - title: Exa docs, Exa MCP\n    url: https://exa.ai/docs/get-started/exa-mcp\n    accessed: 2026-10-02\n  - title: Jina Reader\n    url: https://jina.ai/reader\n    accessed: 2026-10-02\n  - title: OpenAlex Help, Pricing overview\n    url: https://help.openalex.org/access/pricing\n    accessed: 2026-10-02\n  - title: Semantic Scholar API\n    url: https://www.semanticscholar.org/product/api\n    accessed: 2026-10-02\n  - title: arXiv, Terms of Use for arXiv APIs\n    url: https://info.arxiv.org/help/api/tou.html\n    accessed: 2026-10-02\n  - title: Langfuse LICENSE\n    url: https://github.com/langfuse/langfuse/blob/main/LICENSE\n    accessed: 2026-10-02\n  - title: E2B docs, Internet access\n    url: https://docs.e2b.dev/network/internet-access\n    accessed: 2026-10-02\nrelated: [web-search-apis, run-untrusted-code-in-a-sandbox, code-sandboxes, give-an-agent-long-term-memory, connect-an-agent-to-a-remote-mcp-server]\nlast_verified: 2026-10-02\npublished: 2026-10-02\n---\n\n## How the pieces fit\n\nA research agent runs the same loop over and over: plan the question, search, read, take notes, check, write. The stack above maps one component to each step.\n\n1. **Search wide, then deep.** Start with a web search API ([Exa](https://indexagentica.com/entries/exa-api/), with [Tavily](https://indexagentica.com/entries/tavily-api/), [Brave Search](https://indexagentica.com/entries/brave-search-api/) or [Parallel](https://indexagentica.com/entries/parallel-api/) as alternatives) for recent and general sources, and the scholarly APIs ([OpenAlex](https://indexagentica.com/entries/openalex-api/), [Semantic Scholar](https://indexagentica.com/entries/semantic-scholar-api/), [arXiv](https://indexagentica.com/entries/arxiv-api/)) for peer-reviewed work and citation trails. Search results are snippets; don't let the agent cite a snippet.\n2. **Read the source.** Fetch full pages as Markdown with [Firecrawl](https://indexagentica.com/entries/firecrawl-api/) or [Jina Reader](https://indexagentica.com/entries/jina-reader/), and fall back to a real browser ([Playwright MCP](https://indexagentica.com/entries/playwright-mcp/)) when a page needs JavaScript or interaction.\n3. **Compute in a sandbox.** When the question needs numbers, have the agent write analysis code and run it in [E2B](https://indexagentica.com/entries/e2b/), not on your machine. Load the data, then cut the network. The [sandboxing guide](https://indexagentica.com/guides/run-untrusted-code-in-a-sandbox/) explains why.\n4. **Trace everything.** Send every model call and tool call to [Langfuse](https://indexagentica.com/entries/langfuse/). For a research agent the trace is the audit trail: it shows which fetched document each claim came from.\n\nThe harness here is the [Claude Agent SDK](https://indexagentica.com/entries/claude-agent-sdk/); the [OpenAI Agents SDK](https://indexagentica.com/entries/openai-agents-sdk/) or [LangGraph](https://indexagentica.com/entries/langgraph/) work the same way if you prefer them. If your harness speaks MCP, most of these components can be connected as MCP servers instead of custom tools; see the [remote MCP guide](https://indexagentica.com/guides/connect-an-agent-to-a-remote-mcp-server/).\n\n## Practical limits to plan for\n\n- **Rate limits.** arXiv asks for at most one request every three seconds across all your machines. Semantic Scholar's unauthenticated pool is shared and throttled. OpenAlex's free tier is a daily usage budget tied to an API key. Put a queue and a cache in front of the scholarly APIs, and give each its own key.\n- **Prompt injection.** Every page the agent reads is untrusted input. Keep the browser and fetch tools away from credentials, and don't give the same agent a tool that can send email or spend money.\n- **Citation hygiene.** Require a URL or DOI for every claim, and have a final pass re-fetch each cited source and check that the quoted text is there. Search snippets and model memory are not sources.\n- **Memory across sessions.** For long projects, keep notes in files the agent re-reads, or add a memory layer; see [giving an agent long-term memory](https://indexagentica.com/guides/give-an-agent-long-term-memory/).\n\nTo choose between the search APIs, see the [web search APIs comparison](https://indexagentica.com/compare/web-search-apis/).\n",
  "html": "<h2>How the pieces fit</h2>\n<p>A research agent runs the same loop over and over: plan the question, search, read, take notes, check, write. The stack above maps one component to each step.</p>\n<ol><li><strong>Search wide, then deep.</strong> Start with a web search API (<a href=\"/entries/exa-api/\">Exa</a>, with <a href=\"/entries/tavily-api/\">Tavily</a>, <a href=\"/entries/brave-search-api/\">Brave Search</a> or <a href=\"/entries/parallel-api/\">Parallel</a> as alternatives) for recent and general sources, and the scholarly APIs (<a href=\"/entries/openalex-api/\">OpenAlex</a>, <a href=\"/entries/semantic-scholar-api/\">Semantic Scholar</a>, <a href=\"/entries/arxiv-api/\">arXiv</a>) for peer-reviewed work and citation trails. Search results are snippets; don&#39;t let the agent cite a snippet.</li><li><strong>Read the source.</strong> Fetch full pages as Markdown with <a href=\"/entries/firecrawl-api/\">Firecrawl</a> or <a href=\"/entries/jina-reader/\">Jina Reader</a>, and fall back to a real browser (<a href=\"/entries/playwright-mcp/\">Playwright MCP</a>) when a page needs JavaScript or interaction.</li><li><strong>Compute in a sandbox.</strong> When the question needs numbers, have the agent write analysis code and run it in <a href=\"/entries/e2b/\">E2B</a>, not on your machine. Load the data, then cut the network. The <a href=\"/guides/run-untrusted-code-in-a-sandbox/\">sandboxing guide</a> explains why.</li><li><strong>Trace everything.</strong> Send every model call and tool call to <a href=\"/entries/langfuse/\">Langfuse</a>. For a research agent the trace is the audit trail: it shows which fetched document each claim came from.</li></ol>\n<p>The harness here is the <a href=\"/entries/claude-agent-sdk/\">Claude Agent SDK</a>; the <a href=\"/entries/openai-agents-sdk/\">OpenAI Agents SDK</a> or <a href=\"/entries/langgraph/\">LangGraph</a> work the same way if you prefer them. If your harness speaks MCP, most of these components can be connected as MCP servers instead of custom tools; see the <a href=\"/guides/connect-an-agent-to-a-remote-mcp-server/\">remote MCP guide</a>.</p>\n<h2>Practical limits to plan for</h2>\n<ul><li><strong>Rate limits.</strong> arXiv asks for at most one request every three seconds across all your machines. Semantic Scholar&#39;s unauthenticated pool is shared and throttled. OpenAlex&#39;s free tier is a daily usage budget tied to an API key. Put a queue and a cache in front of the scholarly APIs, and give each its own key.</li><li><strong>Prompt injection.</strong> Every page the agent reads is untrusted input. Keep the browser and fetch tools away from credentials, and don&#39;t give the same agent a tool that can send email or spend money.</li><li><strong>Citation hygiene.</strong> Require a URL or DOI for every claim, and have a final pass re-fetch each cited source and check that the quoted text is there. Search snippets and model memory are not sources.</li><li><strong>Memory across sessions.</strong> For long projects, keep notes in files the agent re-reads, or add a memory layer; see <a href=\"/guides/give-an-agent-long-term-memory/\">giving an agent long-term memory</a>.</li></ul>\n<p>To choose between the search APIs, see the <a href=\"/compare/web-search-apis/\">web search APIs comparison</a>.</p>"
}
