# Overview

Edison Scientific is the commercial spinout of [FutureHouse](https://www.futurehouse.org/). We build AI agents that do scientific research: reading papers, analyzing data, forming hypotheses, and producing cited reports. The platform is available today at [platform.edisonscientific.com](https://platform.edisonscientific.com/).

### What the platform does

The Edison Platform is a single environment for scientific R\&D. It brings together literature synthesis, data analysis, molecular design, and novelty assessment through a set of specialized agents. The goal is straightforward: take research tasks that currently take scientists weeks or months and compress them substantially.

### The agents

[**Kosmos**](/agents#kosmos) is our flagship AI Scientist. Given a research objective and one or more datasets, Kosmos autonomously reads literature, writes and executes analysis code, generates hypotheses, and produces a comprehensive cited report. A single run involves reading \~1,500 papers and executing \~42,000 lines of code. Every conclusion is fully auditable. You can trace any finding back to the specific code or literature passage that produced it. Our beta users estimated that a Kosmos run accomplishes work equivalent to roughly 6 months of a PhD or postdoctoral scientist.

[**Literature**](/agents#literature) handles literature search and synthesis. It accesses 175M+ papers, trials, and patents with native understanding of citation graphs, journal quality, and clinical trial data. You can ask it a complex scientific question and get a high-accuracy, cited response, or task it with a deep literature review synthesizing conflicting evidence across hundreds of papers.

[**Analysis**](/agents#analysis) specializes in processing complex experimental data, including flow cytometry, RNA-seq, and other biological datasets. It turns raw data into detailed analyses, statistical results, and publication-ready figures.

[**Precedent**](/agents#precedent) determines whether a research idea has been tried before. It searches across fields to assess novelty and identify gaps, helping you avoid duplicating existing work and focus on what's actually new.

[**Molecules**](/agents#chemistry) is a chemistry-focused agent for molecular design and analysis.

### Access

Edison maintains a generous free tier for academics. Researchers who need higher rate limits or additional features can subscribe to paid plans. If you run into issues or want to learn more, reach out at <contact@edisonscientific.com>.


# Agents

## Kosmos <a href="#kosmos" id="kosmos"></a>

*Your autonomous AI scientist for end-to-end research.*

Kosmos conducts complete research workflows — from reading scientific literature and analyzing data to forming and testing hypotheses and producing comprehensive, citable reports.

#### Key Capabilities

Kosmos can digest over 1,500 research papers and execute more than 42,000 lines of analysis code in a single run. It operates through multiple AI agents working in parallel, sharing information through structured "world models." Every conclusion is fully auditable, so you can trace any finding back to its original code or scientific source. Kosmos also generates publication-ready figures and data visualizations alongside its written analysis.

#### Example Discoveries

* Suggested that high circulating levels of SOD2 protein may reduce myocardial fibrosis in humans
* Identified a possible mechanism linking a specific genetic variant (SNP) to reduced risk of Type 2 diabetes

#### When to Use Kosmos

Kosmos is ideal for complex, high-dimensional datasets (such as scRNA-seq, proteomics, or environmental parameters), comprehensive research projects that require both literature synthesis and data analysis, generating publication-ready reports with validated findings, and discovering novel patterns across multiple datasets.

For more information on how to interact with Kosmos, please see our [guide for best practices](https://sites.gitbook.com/preview/site_5LYpi/guides/best-practices-for-optimizing-kosmos-workflows).

To learn more about Kosmos, please read our blog post [Kosmos: An AI Scientist for Autonomous Discovery](https://edisonscientific.com/articles/announcing-kosmos).

## Literature <a href="#literature" id="literature"></a>

*Fast, accurate, and deeply cited answers to your scientific questions.*

Literature handles search and synthesis tasks with high accuracy, drawing from over 175 million papers, trials, and patents. It natively understands citation graphs, journal quality, clinical trial data, and structured sources.

#### When to Use Literature

Literature is a great fit when you need quick, high-quality answers to complex scientific questions, deep literature reviews and comprehensive synthesis, analysis of conflicting evidence across hundreds of papers, complete reports with abstracts and conclusions, or a guided exploration of a new research domain.

#### Example Use Cases

* *"What are the known genetic associations with pancreatic beta cell dysfunction?"* — Get a quick, precise, and fully cited response.
* *"Analyze contradictory findings on COBLL1's role in adipocyte differentiation"* — Receive a comprehensive synthesis of conflicting evidence.
* Preparing journal club presentations with dozens of synthesized papers.

Literature also supports diagram creation and figure pullouts.

## Analysis <a href="#analysis" id="analysis"></a>

*Turn your experimental data into detailed, actionable insights.*

Analysis specializes in processing complex experimental data — from lab results and RNA-sequencing data to other biological datasets — and transforms them into thorough analyses that answer your research questions.

#### When to Use Analysis

Analysis works well for processing flow cytometry data, RNA-seq analysis, statistical analysis of experimental results, data visualization and figure generation, and multi-modal dataset integration.

#### Example Use Case

After running flow cytometry experiments on GFP-transfected cells with intracellular markers (such as Insulin, NKX6.1, and PDX1), Analysis can process your data, generate visualizations, and identify statistically significant populations or expression patterns.

Analysis also includes protein tools and additional figure generation capabilities.

## Precedent <a href="#precedent" id="precedent"></a>

*Find out if your research idea has been tried before.*

Precedent (formerly known as "Has Anyone") searches across fields to determine whether your hypotheses are truly novel, helping you focus your efforts where they'll have the most impact.

#### When to Use Precedent

Precedent is perfect for answering "Has anyone done this before?", identifying research gaps, avoiding duplication of existing work, and finding innovative approaches in the literature.

#### Example Use Case

Before starting SSR1 knock-in experiments for pseudoislets, Precedent can search the literature to confirm whether similar lentiviral transduction approaches with this specific genetic variant have been attempted, helping you pinpoint what's truly novel about your approach.

## Molecules <a href="#chemistry" id="chemistry"></a>

*A chemistry-focused agent for molecular design and analysis.*

Given a molecule (via SMILES, CAS number, or IUPAC name), it can predict ADMET properties, plan retrosynthesis routes, perform safety assessments, search chemical databases like ChEMBL, and propose optimized lead candidates.

#### When to Use Molecules

Molecules is a good fit for ADMET property prediction, retrosynthesis planning, molecular property calculation (logP, logS, QED, PSA), safety and toxicity assessments, database searches for similar compounds, and iterative lead optimization.

#### Example Use Case

You might ask Molecules to calculate ADMET properties for a given SMILES string, search ChEMBL for similar anti-inflammatory compounds, run safety assessments on the top candidates, and propose modifications to improve solubility.


# EdisonScientific Cookbook

Shared documentation and cookbooks to help accelerate the pace of scientific research.

***

#### Libraries & Frameworks

<table data-view="cards"><thead><tr><th></th><th data-hidden data-card-cover data-type="files"></th></tr></thead><tbody><tr><td><a href="/spaces/M2MYoUWqZzuxV0cP6NbC/pages/J4rBDgx1WnPxb3QqSH1u">PaperQA2 — High accuracy scientific RAG library</a></td><td><a href="/files/B4PbVl50ElY5TymKSKm3">/files/B4PbVl50ElY5TymKSKm3</a></td></tr><tr><td><a href="/spaces/KnMaW0v306jmeMSo6Ygv/pages/rKlxpWcQKT1qBvnO7BQA">Aviary — Environments for LLM agents</a></td><td><a href="/files/M31IR57HEy7KUbQUNCCv">/files/M31IR57HEy7KUbQUNCCv</a></td></tr><tr><td><a href="/spaces/EDDp78urTJPGaSa6xOLf">LDP (Language Decision Processes) — LLM agents library</a></td><td><a href="/files/7yJuoLVzJL87LT0g7A4H">/files/7yJuoLVzJL87LT0g7A4H</a></td></tr></tbody></table>

***

#### Cookbooks & How to guides

<table data-view="cards"><thead><tr><th></th><th data-hidden data-card-cover data-type="image">Cover image</th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><a href="/spaces/M2MYoUWqZzuxV0cP6NbC/pages/U5GlM9v31VNfERZVCK3S">Querying clinical trials with PaperQA2 </a></td><td><a href="/files/y4sApGSmPzZ8OOxQq1xP">/files/y4sApGSmPzZ8OOxQq1xP</a></td><td></td></tr><tr><td><a href="/spaces/M2MYoUWqZzuxV0cP6NbC/pages/8hrctnlc8JAVlPJ4GxP8">Using OpenReview papers with PaperQA2</a></td><td><a href="/files/wrLpwq21NzE8ga5B95Ss">/files/wrLpwq21NzE8ga5B95Ss</a></td><td></td></tr><tr><td><a href="/spaces/LsJxaOaQp0Y25Smq6TL7/pages/Fb9FLHd18RQjLJNQIQ0h">Improving Kosmos queries with our guidelines</a></td><td><a href="https://cdn.prod.website-files.com/69021edc9ce0757fd9564deb/690934278e889fa88a6b1877_Rectangle%20477-p-1600.png">https://cdn.prod.website-files.com/69021edc9ce0757fd9564deb/690934278e889fa88a6b1877_Rectangle%20477-p-1600.png</a></td><td></td></tr><tr><td><a href="/spaces/LsJxaOaQp0Y25Smq6TL7/pages/iXy3ATXAHGxNwXX611oW">Master writing queries for Edison Molecules</a></td><td><a href="/files/0cuno4TByUbmSaDQ55K7">/files/0cuno4TByUbmSaDQ55K7</a></td><td></td></tr><tr><td><a href="/spaces/LsJxaOaQp0Y25Smq6TL7/pages/Dhw5AFysB0ITAuEkWnor">Edison Analysis API tutorial</a></td><td><a href="https://cdn.prod.website-files.com/69021edc9ce0757fd9564deb/691e9cabddc2ff83f1125d2a_ChatGPT%20Image%20Nov%2019%2C%202025%20at%2008_44_08%20PM.png">https://cdn.prod.website-files.com/69021edc9ce0757fd9564deb/691e9cabddc2ff83f1125d2a_ChatGPT%20Image%20Nov%2019%2C%202025%20at%2008_44_08%20PM.png</a></td><td></td></tr></tbody></table>


# Best practices for optimizing Kosmos workflows

Kosmos is a data-driven discovery agent. To get the most out of each Kosmos run, we recommend following these guidelines:

### 1. Provide a clear and feasible research objective that requires iteration.

The research objective can be broad and exploratory or focused and hypothesis-driven. While Kosmos can handle multiple objectives, it performs best with a single, well-defined objective. It should have scope for iterative hypothesis generation and testing. The answer should not be obvious after reading a few papers or conducting a single data analysis. A scientific collaborator in a relevant field should be able to reasonably complete the task within weeks or months.‍

### 2. Provide enough context in the research objective.

The research objective should provide sufficient scientific and experimental context as a starting point for data analysis or literature search. While Kosmos is able to conduct literature search and read all provided datasets, it will greatly benefit from context highlighting nuances unique to your research group, such as key experimental design choices, non-obvious assumptions common in the field, or atypical data-handling protocols. Phrase it as you would explain to an experienced colleague who just joined your team.

Here are some examples of ways to improve research objective prompts:

#### Inappropriate research objectives

* Which of these tissues have a lower ECM pathway expression score?
* List all the genes that are differentially expressed (p < 0.05) between the 'control' and 'drought' RNA-seq samples.
* How is graphene used for water desalination? Suggest cross-linking agents that are compatible with our method.
* Analyze the attached county-level public health dataset. Does ice cream consumption correlate with asthma rates?

#### Good research objectives

* Compare ECM pathway expression levels in the melanoma tissue samples. Propose hypotheses on mechanisms driving these changes and downstream functional consequences.
* Analyze the RNA-seq dataset provided from Arabidopsis thaliana leaves, comparing 'control' (well-watered) to 'drought' (water withholding) samples across a time course. Identify the key regulatory pathways that mediate acclimation. Are there any transcription factor(s) outside the well-known ABA pathway that may be a master regulator of this response?
* We are investigating the use of laminated graphene oxide (GO) membranes for water desalination. A major problem is that these membranes swell and lose selectivity when hydrated. I have attached our recent experimental results showing interlayer spacing vs. salt rejection. Identify the top 3-5 most promising cross-linking agents (e.g., diamines, metal ions) that have been proven to control GO membrane swelling and propose which one would be most compatible with our current layer-by-layer fabrication method.
* The provided county-level public health dataset includes data on disease prevalence, environmental factors, and socio-demographics. Our central hypothesis is that counties with higher air pollution (PM2.5) will have increased asthma prevalence, but we suspect this effect is strongest in low-income communities. Build a regression model to test this hypothesis. Control for confounding variables: this must include population density, average age, and regional differences in pollen, so please treat these as covariates. Propose a follow-up analysis to explore the specific impact of pollen vs. PM2.5.

### 3. Provide sufficient context in the data description.

Ensure that Kosmos has the necessary context about the data provided. For example, if providing a table, ensure all column names are intuitively labeled, or provide an additional sheet that describes what each column name means in more detail. A scientific collaborator in a relevant field should be able to interpret the data with the provided research objective without seeking further clarification.

### 4. Provide complex data.

While Kosmos can operate on simple data (e.g. a csv containing a list of gene names), Kosmos makes the most interesting discoveries when given complex high dimensional data, such as scRNA-seq, proteomics, or environmental parameters across multiple samples or timepoints. You can also provide multiple datasets that are relevant to the research objective. There is no limit on the number of files you provide Kosmos as long as total uncompressed size of your dataset is under 5GB. Kosmos is capable of managing workspaces with 100s of files.

### 5. Provide properly labeled and processed data.

The dataset is used as a starting point for exploratory analysis. While Kosmos can perform quality control steps and is able to correct for artifacts such as batch effects, Kosmos is most powerful when operating on high quality, processed data. We do not recommend getting Kosmos to process your data such as mapping of raw sequencing files or annotating raw imaging data.

### 6. Iterate.

Take some time to write your first Kosmos query so you can avoid the common fallacies listed above. However, trial and error will be your best guide to develop a deep intuition on how to best use the system. Start with a familiar dataset or topic. Since runtime scales with dataset size, begin small for faster results and quicker iteration.


# Best practices for interacting with Edison Molecules

Edison Molecules is a chemistry-focused scientific agent. To get the most out of each Edison Molecules run, follow these guidelines for formulating effective queries:

### 1. Be specific about molecular inputs.

When asking about molecules, provide explicit identifiers such as:

* SMILES strings
* CAS numbers
* IUPAC names

Edison Molecules can handle multiple representations, but being explicit reduces ambiguity. A chemist familiar with your field should be able to unambiguously identify the molecule(s) from your query.

| Insufficient queries                         | More detailed queries                                                                                   |
| -------------------------------------------- | ------------------------------------------------------------------------------------------------------- |
| Tell me about anti-inflammatory drugs.       | What are the ADMET properties for aspirin (SMILES: `CC(=O)OC1=CC=CC=C1C(=O)O` or CAS: 50-78-2)?         |
| What are the functional groups of vitamin C? | List the functional groups available in ascorbic acid (SMILES: `C([C@@H]([C@@H]1C(=C(C(=O)O1)O)O)O)O`). |
| Find compounds to be used as painkillers.    | Search for molecules similar to acetaminophen (SMILES: `CC(=O)NC1=CC=C(C=C1)O`) in the ChEMBL database. |

### 2. Specify desired outputs clearly.

Clearly state what you need from Edison Molecules. It can be a synthesis route, molecular property prediction, safety assessment, literature-backed answer, or a combination of these. The more specific you are about the outputs you need, the better Edison Molecules can select appropriate tools and provide actionable results.

| Insufficient queries                          | More detailed queries                                                                                                                                                                |
| --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Tell me about `CN1C=NC2=C1C(=O)N(C(=O)N2C)C`. | What are the ADMET properties for `CN1C=NC2=C1C(=O)N(C(=O)N2C)C`? Also, what are its GHS classification and LD50 value?                                                              |
| Is CC(=O)OC1=CC=CC=C1C(=O)O safe?             | Perform a safety assessment for the molecule with SMILES `CC(=O)OC1=CC=CC=C1C(=O)O`, including GHS classification, LD50 value, chemical weapons screening, and toxicity predictions. |
| Can you help me make aspirin?                 | I need a synthesis route, ADMET property predictions, and a safety assessment for the target molecule with SMILES `CC(=O)OC1=CC=CC=C1C(=O)O`.                                        |

### 3. Use proper chemical terminology and notation.

Employ standard chemical nomenclature and notation. For example, use SMILES representation for reactions: `reactants>reagents>products`. Use standard property names (e.g., ADMET properties) when requesting specific molecular properties. This helps Edison Molecules understand your intent and select the most appropriate computational tools.

| Insufficient queries                           | More detailed queries                                                                                                                                                                                                    |
| ---------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| What happens if I mix ethanol and acetic acid? | Predict the product of the reaction with reaction SMILES: `CCO.CC(=O)O>>CC(=O)OC`. Calculate the reaction enthalpy.                                                                                                      |
| How do I make an ester from alcohol and acid?  | Design a synthesis route for ethyl acetate using reaction SMILES: `CCO.CC(=O)O>>CC(=O)OC` with appropriate catalysts and conditions.                                                                                     |
| Check the drug properties of quercetin.        | Calculate ADMET properties (specifically human intestinal absorption, blood-brain barrier permeability, and cytochrome P450 inhibition) for the molecule with SMILES `C1=CC(=C(C=C1C2=C(C(=O)C3=C(C=C(C=C3O2)O)O)O)O)O`. |

### 4. Break down complex queries.

Structure multi-part questions logically so Edison Molecules can create an effective execution plan. While Edison Molecules can handle multi-step queries and longer workflows, clearly organizing your query helps ensure all components are addressed systematically.

| Insufficient queries                                                  | More detailed queries                                                                                                                                                                                                                                                         |
| --------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Find similar drugs to aspirin.                                        | (1) Analyze aspirin (SMILES: `CC(=O)OC1=CC=CC=C1C(=O)O`) for ADMET properties. (2) Search ChEMBL for similar anti-inflammatory compounds. (3) Perform safety assessments on the top 5 candidates. (4) Propose modifications to improve solubility while maintaining efficacy. |
| I need everything about caffeine and how to make it and what it does. | I need you to give me information about caffeine. Follow these steps: (1) Calculate molecular properties (logP, logS, QED) for SMILES `CN1C=NC2=C1C(=O)N(C(=O)N2C)C`. (2) Design a retrosynthesis route. (3) Predict biological activity targets.                             |
| Find drugs for diabetes, check if they work, and make new ones.       | (1) Search ChEMBL for approved diabetes drugs targeting INSR. (2) Analyze their binding affinities and ADMET properties. (3) Propose 5 novel small molecule candidates with improved properties.                                                                              |

### 5. Provide context when relevant.

Include background information about your use case (e.g., "for drug development" or "for a research synthesis") to help Edison Molecules select appropriate tools and safety considerations. Context about your constraints, goals, or specific requirements enables Edison Molecules to provide more targeted and useful responses.

| Insufficient queries                                  | More detailed queries                                                                                                                                                                                   |
| ----------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Find similar molecules to `CC(=O)OC1=CC=CC=C1C(=O)O`. | Search the ChEMBL database for molecules similar to aspirin (SMILES: `CC(=O)OC1=CC=CC=C1C(=O)O`) for drug repurposing. Return the top 10 candidates with their development phases and bioactivity data. |

### 6. Request specific properties or analyses.

Instead of asking vaguely about a molecule, specify what you need. For example, request specific ADMET properties, synthetic accessibility scores, solubility predictions, or toxicity data. This allows Edison Molecules to use the most appropriate computational tools and provide quantitative, actionable results.

| Insufficient queries                      | More detailed queries                                                                                                                                                                                                       |
| ----------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Modify quercetin to make it more soluble. | Suggest three different substitution modifications to quercetin to make its aqueous solubility (logS) higher. Show me the suggested molecules in your final answer and base your answer in predicted solubility data.       |
| Is aspirin drug-like?                     | Calculate the drug-likeness score (QED), synthetic accessibility score (SAscore), and Lipinski's Rule of Five violations for `CC(=O)OC1=CC=CC=C1C(=O)O`.                                                                    |
| What are the properties of caffeine?      | Calculate the following properties for caffeine (SMILES: `CN1C=NC2=C1C(=O)N(C(=O)N2C)C`): (1) aqueous solubility (logS), (2) partition coefficient (logP), (3) polar surface area (PSA), and (4) number of rotatable bonds. |

### 7. Ask actionable questions that leverage Edison Molecules' toolset.

Frame queries that can be answered using Edison Molecules' computational capabilities rather than pure theoretical discussions without computational support. Edison Molecules excels at property prediction, synthesis planning, reaction analysis, database searches, and literature-enhanced discovery.

| Insufficient queries                              | More detailed queries                                                                                                                                                          |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| How do I make aspirin?                            | Design a retrosynthesis route for the target molecule with SMILES `CC(=O)OC1=CC=CC=C1C(=O)O`. Please identify starting materials and propose reaction steps.                   |
| Can you synthesize caffeine?                      | Propose a synthesis route for caffeine (SMILES: `CN1C=NC2=C1C(=O)N(C(=O)N2C)C`), including retrosynthetic analysis, reactants pricing, and reaction conditions where possible. |
| What drugs treat diabetes?                        | Propose small molecule binders for the insulin receptor (INSR gene symbol). Propose up to 10 candidates and analyze their drug-likeness using QED scores.                      |
| Explain how this ethanol reacts with acetic acid. | Predict the product of the reaction with SMILES `CCO.CC(=O)O>>`, calculate the reaction enthalpy, identify the mechanism, and suggest optimal catalysts and conditions.        |

### 8. Iterate.

Take some time to write your first Edison Molecules query so you can avoid the common pitfalls listed above. Starting with simpler queries will give you faster results and allow quicker iteration, helping you learn how to use Edison Molecules more effectively and get the most out of the system. Your first query does not have to be perfect. Use Edison Molecules interactively to refine your query:

1. Start simple
   * Begin with a well-known molecule or reaction and a small set of properties or a basic retrosynthesis.
2. Inspect the outputs
   * Check if the properties, synthesis steps, or hits match your expectations.
3. Refine your query
   * If you need more detail, add explicit properties, constraints, or additional steps.
   * Example: "Now also calculate logS and propose modifications to improve solubility."
4. Repeat
   * Use the results of one Edison Molecules run as input for the next (e.g., take top hits and ask for safety assessments or optimization ideas).
   * Edison Molecules also accepts follow-up questions to the previous run.


# Quickstart

Documentation and tutorials for programmatically interacting with the Edison platform.

## Installation <a href="#installation" id="installation"></a>

```bash
uv pip install edison-client
```

## Authentication <a href="#authentication" id="authentication"></a>

In order to use the `EdisonClient`, you need to authenticate yourself. Authentication is done by providing an API key, which can be obtained directly from your [profile page in the Edison platform](https://platform.edisonscientific.com/profile).

To access your API key:

1. Create an account on the [Edison platform](https://platform.edisonscientific.com/).
2. Click on the "Account" button on the left sidebar. Then click on the "Profile" button.
3. Under "API Tokens", click "Create New Token."

## Basic task running <a href="#basic-task-running" id="basic-task-running"></a>

```python
from edison_client import EdisonClient, JobNames

client = EdisonClient(
    api_key="<your_api_key>",
)

task_data = {
    "name": JobNames.LITERATURE,
    "query": "Which neglected diseases had a treatment developed by artificial intelligence?",
}

task_response = client.run_tasks_until_done(task_data)
```


# Tasks

Run and manage tasks with the Edison client.

## Overview <a href="#overview" id="overview"></a>

Edison client implements a RestClient (called `EdisonClient`) with the following functionalities:

* [Simple task running](https://edisonscientific.gitbook.io/edison-cookbook/edison-client#simple-task-running): `run_tasks_until_done(TaskRequest)` or `await arun_tasks_until_done(TaskRequest)`
* [Asynchronous tasks](https://edisonscientific.gitbook.io/edison-cookbook/edison-client#asynchronous-tasks): `get_task(task_id)` or `aget_task(task_id)` and `create_task(TaskRequest)` or `acreate_task(TaskRequest)`

To create a `EdisonClient`, you need to pass an Edison Scientific platform api key (see [Authentication](/edison-client#authentication)):

## Task types <a href="#task-types" id="task-types"></a>

In the Edison platform, we define the deployed combination of an agent and an environment as a `job`.

To invoke a job, we need to submit a `task` (also called a `query`) to it. `EdisonClient` can be used to submit tasks/queries to available jobs in the Edison platform.

Using an `EdisonClient` instance, you can submit tasks to the platform by calling the `create_task` method, which receives a `TaskRequest` (or a dictionary with `kwargs`) and returns the task ID.

Aiming to make the submission of tasks as simple as possible, we have created a `JobNames` `enum` that contains the available task types.

Please note that Kosmos is not available via API.

| Alias                      | Task type         | Description                                                                                                                                              |
| -------------------------- | ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `JobNames.LITERATURE`      | Literature search | Ask a question of scientific data sources, and receive a high-accuracy, cited response. Built with [PaperQA3](https://github.com/Future-House/paper-qa). |
| `JobNames.LITERATURE_HIGH` | Literature search | Ask a question of scientific data sources, and receive a high-accuracy, cited response. High reasoning mode enabled for SOTA performance.                |
| `JobNames.ANALYSIS`        | Data analysis     | Turn biological datasets into detailed analyses answering your research questions.                                                                       |
| `JobNames.PRECEDENT`       | Precedent search  | Formerly known as HasAnyone, query if anyone has ever done something in science.                                                                         |
| `JobNames.MOLECULES`       | Chemistry tasks   | A new iteration of ChemCrow, Phoenix uses cheminformatics tools to do chemistry. Good for planning synthesis and designing new molecules.                |

## Submitting tasks <a href="#submitting-tasks" id="submitting-tasks"></a>

Using `JobNames`, the task submission looks like this:

```python
from edison_client import EdisonClient, JobNames

client = EdisonClient(
    api_key="<your_api_key>",
)

task_data = {
    "name": JobNames.PRECEDENT,
    "query": "Has anyone tested therapeutic exerkines in humans or NHPs?",
}

task_response = client.run_tasks_until_done(task_data)

print(task_response)
```

## Asynchronous tasks <a href="#asynchronous-tasks" id="asynchronous-tasks"></a>

Sometimes you may want to submit many jobs, while querying results at a later time. The platform API supports this, as shown below.

```python
import asyncio
from edison_client import EdisonClient, JobNames


async def main():
    client = EdisonClient(
        api_key="<your_api_key>",
    )

    task_data = {
        "name": JobNames.PRECEDENT,
        "query": "Has anyone tested therapeutic exerkines in humans or NHPs?",
    }

    task_response = await client.arun_tasks_until_done(task_data)
    return task_response


# For Python 3.7+
if __name__ == "__main__":
    task_response = asyncio.run(main())
```

## Batch task submission <a href="#batch-task-submission" id="batch-task-submission"></a>

In either the sync or the async code, collections of tasks can be given to the client to run them in a batch:

```python
import asyncio
from edison_client import EdisonClient, JobNames


async def main():
    client = EdisonClient(
        api_key="<your_api_key>",
    )

    task_data = [{
        "name": JobNames.PRECEDENT,
        "query": "Has anyone tested therapeutic exerkines in humans or NHPs?",
    },
    {
        "name": JobNames.LITERATURE,
        "query": "Are there any clinically validated therapeutic exerkines for humans?",
    }
    ]

    task_responses = await client.arun_tasks_until_done(task_data)
    return task_responses


# For Python 3.7+
if __name__ == "__main__":
    task_responses = asyncio.run(main())
```

## Task continuation <a href="#task-continuation" id="task-continuation"></a>

Once a task is submitted and the answer is returned, Edison platform allow you to ask follow-up questions to the previous task. It is also possible through the platform API. To accomplish that, we can use the `runtime_config` we discussed in the [Simple task running](https://edisonscientific.gitbook.io/edison-cookbook/edison-client#simple-task-running) section.

```python
from edison_client import EdisonClient, JobNames

client = EdisonClient(
    api_key="<your_api_key>",
)

task_data = {"name": JobNames.LITERATURE, "query": "How many species of birds are there?"}

task_id = client.create_task(task_data)

continued_task_data = {
    "name": JobNames.LITERATURE,
    "query": "From the previous answer, specifically, how many species of crows are there?",
    "runtime_config": {"continued_job_id": task_id},
}

task_result = client.run_tasks_until_done(continued_task_data)
```


# File management

## Upload <a href="#upload" id="upload"></a>

`Edison Analysis` is designed to run data analysis on files provided by the user or caller. To provide `Edison Analysis` with this data, you'll need to upload it to the Edison data storage service. This service is your one stop shop for sharing, storing and updating data to be used in the Edison ecosystem.

To upload a single file to the data storage service:

```python
single_file_upload_response = await client.astore_file_content(
    name="Demo file entry for a single file",
    file_path="./datasets/brain_size_data.csv",  # ADD DATASET PATH HERE
    description="This is a test file that will be be analysed by Edison Analysis",
)
```

To upload a directory to the data storage service:

```python
directory_upload_response = await client.astore_file_content(
    name="Demo file entry for a whole directory",
    file_path="./datasets",  # ADD DATASET FOLDER PATH HERE
    description="This is a directory that will be be analysed by Edison Analysis",
    as_collection=True,
)
```

## Download <a href="#download" id="download"></a>

While the task is executing it will create some artifacts. First the notebook which is where the analysis code will be written and any other artifacts creating during the task.

Once the task has completed you may want to check the contents of the notebook or look through the artifacts generated. To obtain these artifacts, you will need to inspect the output of the agent's final `environment_frame`

```python
output_data = job_result.environment_frame["state"]["info"]["output_data"]
print(output_data)
```

```python
for output_file in output_data:
    download_response = await client.afetch_data_from_storage(
        data_storage_id=output_file["entry_id"]
    )

    # Note there are two potential outcomes here. One where the client downloads
    # the file to your local filesystem if it's above ~10MB. The second is where
    # it will return a RawFetchResponse object which contains the raw content.
    print(download_response)
```


# Edison Platform API Documentation

[![PyPI version](https://badge.fury.io/py/edison-client.svg)](https://badge.fury.io/py/edison-client) ![License](https://img.shields.io/badge/License-Apache_2.0-blue.svg) ![PyPI Python Versions](https://img.shields.io/pypi/pyversions/edison-client)

Documentation and tutorials for edison-client, a client for interacting with endpoints of the Edison platform.

* [Installation](#installation)
* [Quickstart](#quickstart)
* [Functionalities](#functionalities)
* [Authentication](#authentication)
* [Simple task running](#simple-task-running)
* [Task Continuation](#task-continuation)
* [Asynchronous tasks](#asynchronous-tasks)

## Installation

```bash
uv pip install edison-client
```

## Quickstart

```python
from edison_client import EdisonClient, JobNames

client = EdisonClient(
    api_key="your_api_key",
)

task_data = {
    "name": JobNames.LITERATURE,
    "query": "Which neglected diseases had a treatment developed by artificial intelligence?",
}

task_response = client.run_tasks_until_done(task_data)
```

A quickstart example can be found in the [client\_notebook.ipynb](https://github.com/Future-House/crow-client-docs/blob/main/docs/client_notebook.ipynb) file, where we show how to submit and retrieve a task, pass runtime configuration to the agent, and ask follow-up questions to the previous task.

## Functionalities

Edison client implements a RestClient (called `EdisonClient`) with the following functionalities:

* [Simple task running](#simple-task-running): `run_tasks_until_done(TaskRequest)` or `await arun_tasks_until_done(TaskRequest)`
* [Asynchronous tasks](#asynchronous-tasks): `get_task(task_id)` or `aget_task(task_id)` and `create_task(TaskRequest)` or `acreate_task(TaskRequest)`

To create a `EdisonClient`, you need to pass an Edison Scientific platform api key (see [Authentication](#authentication)):

```python
from edison_client import EdisonClient

client = EdisonClient(
    api_key="your_api_key",
)
```

## Authentication

In order to use the `EdisonClient`, you need to authenticate yourself. Authentication is done by providing an API key, which can be obtained directly from your [profile page in the Edison platform](https://platform.edisonscientific.com/profile).

## Simple task running

In the Edison platform, we define the deployed combination of an agent and an environment as a `job`. To invoke a job, we need to submit a `task` (also called a `query`) to it. `EdisonClient` can be used to submit tasks/queries to available jobs in the Edison platform. Using a `EdisonClient` instance, you can submit tasks to the platform by calling the `create_task` method, which receives a `TaskRequest` (or a dictionary with `kwargs`) and returns the task id. Aiming to make the submission of tasks as simple as possible, we have created a `JobNames` `enum` that contains the available task types.

The available supported jobs are:

| Alias                      | Available Aliases                                         | Task type         | Description                                                                                                                                              |
| -------------------------- | --------------------------------------------------------- | ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `JobNames.LITERATURE`      | `literature-20260216`, `JobNames.CROW`, `JobNames.FALCON` | Literature Search | Ask a question of scientific data sources, and receive a high-accuracy, cited response. Built with [PaperQA3](https://github.com/Future-House/paper-qa). |
| `JobNames.LITERATURE_HIGH` | `literature-high-20260216`                                | Literature Search | Ask a question of scientific data sources, and receive a high-accuracy, cited response. High reasoning mode enabled for SOTA performance.                |
| `JobNames.ANALYSIS`        | `JobNames.FINCH`                                          | Data Analysis     | Turn biological datasets into detailed analyses answering your research questions.                                                                       |
| `JobNames.PRECEDENT`       | `JobNames.OWL`                                            | Precedent Search  | Formerly known as HasAnyone, query if anyone has ever done something in science.                                                                         |
| `JobNames.MOLECULES`       | `JobNames.PHOENIX`                                        | Chemistry Tasks   | A new iteration of ChemCrow, Phoenix uses cheminformatics tools to do chemistry. Good for planning synthesis and designing new molecules.                |

Using `JobNames`, the task submission looks like this:

```python
from edison_client import EdisonClient, JobNames

client = EdisonClient(
    api_key="your_api_key",
)

task_data = {
    "name": JobNames.PRECEDENT,
    "query": "Has anyone tested therapeutic exerkines in humans or NHPs?",
}

task_response = client.run_tasks_until_done(task_data)

print(task_response.answer)
```

Or if running async code:

```python
import asyncio
from edison_client import EdisonClient, JobNames


async def main():
    client = EdisonClient(
        api_key="your_api_key",
    )

    task_data = {
        "name": JobNames.PRECEDENT,
        "query": "Has anyone tested therapeutic exerkines in humans or NHPs?",
    }

    task_response = await client.arun_tasks_until_done(task_data)
    print(task_response.answer)
    return task_id


# For Python 3.7+
if __name__ == "__main__":
    task_id = asyncio.run(main())
```

Note that in either the sync or the async code, collections of tasks can be given to the client to run them in a batch:

```python
import asyncio
from edison_client import EdisonClient, JobNames


async def main():
    client = EdisonClient(
        api_key="your_api_key",
    )

    task_data = [{
        "name": JobNames.PRECEDENT,
        "query": "Has anyone tested therapeutic exerkines in humans or NHPs?",
    },
    {
        "name": JobNames.LITERATURE,
        "query": "Are there any clinically validated therapeutic exerkines for humans?",
    }
    ]

    task_responses = await client.arun_tasks_until_done(task_data)
    print(task_responses[0].answer)
    print(task_responses[1].answer)
    return task_id


# For Python 3.7+
if __name__ == "__main__":
    task_id = asyncio.run(main())
```

`TaskRequest` can also be used to submit jobs and it has the following fields:

| Field           | Type          | Description                                                                                                               |
| --------------- | ------------- | ------------------------------------------------------------------------------------------------------------------------- |
| id              | UUID          | Optional job identifier. A UUID will be generated if not provided                                                         |
| name            | str           | Name of the job to execute eg. `job-futurehouse-paperqa2`, or using the `JobNames` for convenience: `JobNames.LITERATURE` |
| query           | str           | Query or task to be executed by the job                                                                                   |
| runtime\_config | RuntimeConfig | Optional runtime parameters for the job                                                                                   |

`runtime_config` can receive a `AgentConfig` object with the desired kwargs. Check the available `AgentConfig` fields in the [LDP documentation](https://github.com/Future-House/ldp/blob/main/src/ldp/agent/agent.py#L87). Besides the `AgentConfig` object, we can also pass `timeout` and `max_steps` to limit the execution time and the number of steps the agent can take.

```python
from edison_client import EdisonClient, JobNames
from edison_client.models.app import TaskRequest

client = EdisonClient(
    api_key="your_api_key",
)

task_response = client.run_tasks_until_done(
    TaskRequest(
        name=JobNames.PRECEDENT,
        query="Has anyone tested therapeutic exerkines in humans or NHPs?",
    )
)

print(task_response.answer)
```

A `TaskResponse` will be returned from using our agents. For `LITERATURE` and `PRECEDENT`, we default to a subclass, `PQATaskResponse` which has some key attributes:

| Field                   | Type | Description                                                                     |
| ----------------------- | ---- | ------------------------------------------------------------------------------- |
| answer                  | str  | Answer to your query.                                                           |
| formatted\_answer       | str  | Specially formatted answer with references.                                     |
| has\_successful\_answer | bool | Flag for whether the agent was able to find a good answer to your query or not. |

If using the `verbose` setting, much more data can be pulled down from your `TaskResponse`, which will exist across all agents.

```python
from edison_client import EdisonClient, JobNames
from edison_client.models.app import TaskRequest

client = EdisonClient(
    api_key="your_api_key",
)

task_response = client.run_tasks_until_done(
    TaskRequest(
        name=JobNames.PRECEDENT,
        query="Has anyone tested therapeutic exerkines in humans or NHPs?",
    ),
    verbose=True,
)

print(task_response.environment_frame)
```

In that case, a `TaskResponseVerbose` will have the following fields:

| Field              | Type | Description                                                                                                            |
| ------------------ | ---- | ---------------------------------------------------------------------------------------------------------------------- |
| agent\_state       | dict | Large object with all agent states during the progress of your task.                                                   |
| environment\_frame | dict | Large nested object with all environment data, for PQA environments it includes contexts, paper metadata, and answers. |
| metadata           | dict | Extra metadata about your query.                                                                                       |

## Task Continuation

Once a task is submitted and the answer is returned, Edison platform allow you to ask follow-up questions to the previous task. It is also possible through the platform API. To accomplish that, we can use the `runtime_config` we discussed in the [Simple task running](#simple-task-running) section.

```python
from edison_client import EdisonClient, JobNames

client = EdisonClient(
    api_key="your_api_key",
)

task_data = {"name": JobNames.LITERATURE, "query": "How many species of birds are there?"}

task_id = client.create_task(task_data)

continued_task_data = {
    "name": JobNames.LITERATURE,
    "query": "From the previous answer, specifically, how many species of crows are there?",
    "runtime_config": {"continued_job_id": task_id},
}

task_result = client.run_tasks_until_done(continued_task_data)
```

## Asynchronous tasks

Sometimes you may want to submit many jobs, while querying results at a later time. In this way you can do other things while waiting for a response. The platform API supports this as well rather than waiting for a result.

```python
from edison_client import EdisonClient

client = EdisonClient(
    api_key="your_api_key",
)

task_data = {"name": JobNames.LITERATURE, "query": "How many species of birds are there?"}

task_id = client.create_task(task_data)

# move on to do other things

task_status = client.get_task(task_id)
```

`task_status` contains information about the task. For instance, its `status`, `task`, `environment_name` and `agent_name`, and other fields specific to the job. You can continually query the status until it's `success` before moving on.


# Methods

Every `EdisonClient` method, grouped by area. Each group lists the synchronous methods and their `a`-prefixed asynchronous twins. Async twins take the same arguments and return the same types — just `await` them.

Anywhere a model such as `TaskRequest` is accepted, you can also pass a plain `dict` with the same fields — it's validated into the model for you. Both argument payloads and return values are shown below as example JSON; sample values are placeholders.

### Tasks

#### Sync

**`run_tasks_until_done()`**

Run multiple tasks and wait for them to complete.

```python
run_tasks_until_done(
    task_data: TaskRequest | Collection[TaskRequest],
    verbose: bool = False,
    progress_bar: bool = False,
    timeout: float | None = 2400,
    files: list[str] | None = None,
) -> list[LiteTaskResponse | TaskResponse | TaskResponseVerbose]
```

**Arguments**

```json
{
  "task_data": {
    "task_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "persona_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "name": "string",
    "query": "string",
    "runtime_config": {
      "timeout": 0,
      "max_steps": 0,
      "agent": null,
      "environment_config": {},
      "continued_job_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "world_model_id": "019cd179-8d61-7f65-a79c-b965dda9eac3"
    },
    "caller_target_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "caller_project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "tags": [
      "string"
    ],
    "workspace_id": "string"
  },
  "verbose": false,
  "progress_bar": false,
  "timeout": 2400,
  "files": null
}
```

**Returns** `list[LiteTaskResponse | TaskResponse | TaskResponseVerbose]` — each element is one of:

**`LiteTaskResponse`**

```json
{
  "task_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
  "query": "string",
  "status": "string"
}
```

**`TaskResponse`**

```json
{
  "status": "string",
  "query": "string",
  "user": "string",
  "created_at": "2026-06-03T08:00:00Z",
  "job_name": "string",
  "share_status": "string",
  "permitted_accessors": {},
  "build_owner": "string",
  "environment_name": "string",
  "agent_name": "string",
  "task_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
  "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3"
}
```

**`TaskResponseVerbose`**

```json
{
  "status": "string",
  "query": "string",
  "user": "string",
  "created_at": "2026-06-03T08:00:00Z",
  "job_name": "string",
  "share_status": "string",
  "permitted_accessors": {},
  "build_owner": "string",
  "environment_name": "string",
  "agent_name": "string",
  "task_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
  "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
  "agent_state": [
    {}
  ],
  "environment_frame": {},
  "metadata": {},
  "deployment_config": {
    "commit_hash": "string",
    "git_patch": "string",
    "branch_name": "string",
    "git_status": "string"
  }
}
```

**`create_task()`**

Create a new futurehouse task.

```python
create_task(
    task_data: TaskRequest,
    files: list[str] | None = None,
)
```

**Arguments**

```json
{
  "task_data": {
    "task_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "persona_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "name": "string",
    "query": "string",
    "runtime_config": {
      "timeout": 0,
      "max_steps": 0,
      "agent": null,
      "environment_config": {},
      "continued_job_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "world_model_id": "019cd179-8d61-7f65-a79c-b965dda9eac3"
    },
    "caller_target_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "caller_project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "tags": [
      "string"
    ],
    "workspace_id": "string"
  },
  "files": null
}
```

**Returns** `str` — The new task's `trajectory_id`.

```json
"019cd179-8d61-7f65-a79c-b965dda9eac3"
```

**`get_task()`**

Get details for a specific task.

```python
get_task(
    task_id: str | None = None,
    history: bool = False,
    verbose: bool = False,
    lite: bool = False,
) -> TaskResponse | TaskResponseVerbose | LiteTaskResponse
```

**Returns** `TaskResponse | TaskResponseVerbose | LiteTaskResponse` — one of:

**`TaskResponse`**

```json
{
  "status": "string",
  "query": "string",
  "user": "string",
  "created_at": "2026-06-03T08:00:00Z",
  "job_name": "string",
  "share_status": "string",
  "permitted_accessors": {},
  "build_owner": "string",
  "environment_name": "string",
  "agent_name": "string",
  "task_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
  "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3"
}
```

**`TaskResponseVerbose`**

```json
{
  "status": "string",
  "query": "string",
  "user": "string",
  "created_at": "2026-06-03T08:00:00Z",
  "job_name": "string",
  "share_status": "string",
  "permitted_accessors": {},
  "build_owner": "string",
  "environment_name": "string",
  "agent_name": "string",
  "task_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
  "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
  "agent_state": [
    {}
  ],
  "environment_frame": {},
  "metadata": {},
  "deployment_config": {
    "commit_hash": "string",
    "git_patch": "string",
    "branch_name": "string",
    "git_status": "string"
  }
}
```

**`LiteTaskResponse`**

```json
{
  "task_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
  "query": "string",
  "status": "string"
}
```

**`get_tasks()`**

Fetches trajectories with applied filtering.

```python
get_tasks(
    query_params: TrajectoryQueryParams | None = None,
    project_id: UUID | None = None,
    name: str | None = None,
    user: str | None = None,
    limit: int = 50,
    offset: int = 0,
    sort_by: str = 'created_at',
    sort_order: str = 'desc',
) -> list[dict[str, Any]]
```

**Arguments**

```json
{
  "query_params": null,
  "project_id": null,
  "name": null,
  "user": null,
  "limit": 50,
  "offset": 0,
  "sort_by": "created_at",
  "sort_order": "desc"
}
```

**Returns** `list[dict[str, Any]]`

*Raw, unvalidated dicts (no Pydantic model).*

**`cancel_task()`**

Cancel a specific task/trajectory.

```python
cancel_task(
    task_id: str | None = None,
) -> bool
```

**Returns** `bool`

**`delete_trajectory()`**

Delete a trajectory (soft delete - marks as hidden).

```python
delete_trajectory(
    trajectory_id: UUID,
) -> None
```

**Returns** nothing (`None`).

#### Async

**`arun_tasks_until_done()`**

Awaitable twin of `run_tasks_until_done()` — identical arguments and return type.

```python
arun_tasks_until_done(
    task_data: TaskRequest | Collection[TaskRequest],
    verbose: bool = False,
    progress_bar: bool = False,
    concurrency: int = 10,
    timeout: float | None = 2400,
    files: list[str] | None = None,
) -> list[LiteTaskResponse | TaskResponse | TaskResponseVerbose]
```

**`acreate_task()`**

Awaitable twin of `create_task()` — identical arguments and return type.

```python
acreate_task(
    task_data: TaskRequest,
    files: list[str] | None = None,
)
```

**`aget_task()`**

Awaitable twin of `get_task()` — identical arguments and return type.

```python
aget_task(
    task_id: str | None = None,
    history: bool = False,
    verbose: bool = False,
    lite: bool = False,
) -> TaskResponse | TaskResponseVerbose | LiteTaskResponse
```

**`aget_tasks()`**

Awaitable twin of `get_tasks()` — identical arguments and return type.

```python
aget_tasks(
    query_params: TrajectoryQueryParams | None = None,
    project_id: UUID | None = None,
    name: str | None = None,
    user: str | None = None,
    limit: int = 50,
    offset: int = 0,
    sort_by: str = 'created_at',
    sort_order: str = 'desc',
) -> list[dict[str, Any]]
```

**`adelete_trajectory()`**

Awaitable twin of `delete_trajectory()` — identical arguments and return type.

```python
adelete_trajectory(
    trajectory_id: UUID,
) -> None
```

### Data Storage

#### Sync

**`store_file_content()`**

Store file or directory content in the data storage system.

```python
store_file_content(
    name: str,
    file_path: str | Path,
    description: str | None = None,
    file_path_override: str | Path | None = None,
    as_collection: bool = False,
    manifest_filename: str | None = None,
    ignore_patterns: list[str] | None = None,
    ignore_filename: str = '.gitignore',
    project_id: UUID | None = None,
    dataset_id: UUID | None = None,
    metadata: dict[str, Any] | None = None,
    tags: list[str] | None = None,
    parent_id: UUID | None = None,
    embedding: list[float] | None = None,
) -> DataStorageResponse
```

**Returns** `DataStorageResponse`

```json
{
  "data_storage": {
    "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "name": "string",
    "description": "string",
    "content": "string",
    "status": "pending",
    "embedding": [
      0.0
    ],
    "is_collection": false,
    "tags": [
      "string"
    ],
    "parent_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "dataset_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "file_path": "string",
    "bigquery_schema": null,
    "user_id": "string",
    "created_at": "2026-06-03T08:00:00Z",
    "modified_at": "2026-06-03T08:00:00Z",
    "share_status": "private",
    "short_alias": "string"
  },
  "storage_locations": [
    {
      "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "data_storage_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "storage_config": {
        "storage_type": "string",
        "content_type": "string",
        "content_schema": null,
        "metadata": null,
        "location": "string",
        "signed_url": "string"
      },
      "created_at": "2026-06-03T08:00:00Z"
    }
  ]
}
```

**`store_text_content()`**

Store content as a string in the data storage system.

```python
store_text_content(
    name: str,
    content: str,
    description: str | None = None,
    file_path: str | None = None,
    project_id: UUID | None = None,
    metadata: dict[str, Any] | None = None,
    tags: list[str] | None = None,
    dataset_id: UUID | None = None,
    parent_id: UUID | None = None,
    embedding: list[float] | None = None,
) -> DataStorageResponse
```

**Returns** `DataStorageResponse`

```json
{
  "data_storage": {
    "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "name": "string",
    "description": "string",
    "content": "string",
    "status": "pending",
    "embedding": [
      0.0
    ],
    "is_collection": false,
    "tags": [
      "string"
    ],
    "parent_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "dataset_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "file_path": "string",
    "bigquery_schema": null,
    "user_id": "string",
    "created_at": "2026-06-03T08:00:00Z",
    "modified_at": "2026-06-03T08:00:00Z",
    "share_status": "private",
    "short_alias": "string"
  },
  "storage_locations": [
    {
      "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "data_storage_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "storage_config": {
        "storage_type": "string",
        "content_type": "string",
        "content_schema": null,
        "metadata": null,
        "location": "string",
        "signed_url": "string"
      },
      "created_at": "2026-06-03T08:00:00Z"
    }
  ]
}
```

**`store_link()`**

Store a link/URL in the data storage system.

```python
store_link(
    name: str,
    url: HttpUrl,
    description: str,
    instructions: str,
    api_key: str | None = None,
    metadata: dict[str, Any] | None = None,
    dataset_id: UUID | None = None,
    project_id: UUID | None = None,
    tags: list[str] | None = None,
    parent_id: UUID | None = None,
) -> DataStorageResponse
```

**Returns** `DataStorageResponse`

```json
{
  "data_storage": {
    "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "name": "string",
    "description": "string",
    "content": "string",
    "status": "pending",
    "embedding": [
      0.0
    ],
    "is_collection": false,
    "tags": [
      "string"
    ],
    "parent_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "dataset_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "file_path": "string",
    "bigquery_schema": null,
    "user_id": "string",
    "created_at": "2026-06-03T08:00:00Z",
    "modified_at": "2026-06-03T08:00:00Z",
    "share_status": "private",
    "short_alias": "string"
  },
  "storage_locations": [
    {
      "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "data_storage_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "storage_config": {
        "storage_type": "string",
        "content_type": "string",
        "content_schema": null,
        "metadata": null,
        "location": "string",
        "signed_url": "string"
      },
      "created_at": "2026-06-03T08:00:00Z"
    }
  ]
}
```

**`fetch_data_from_storage()`**

Fetch data from the storage system.

```python
fetch_data_from_storage(
    data_storage_id: UUID | str | None = None,
) -> RawFetchResponse | Path | list[Path] | None
```

**Returns** `RawFetchResponse | Path | list[Path] | None` — one of:

**`RawFetchResponse`**

```json
{
  "filename": "/local/path/to/file",
  "content": "string",
  "entry_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
  "entry_name": "string"
}
```

**`Path`** — `"/local/path/to/file"`

**`list[Path]`** — `["/local/path/to/file"]`

**`None`** — `null` (e.g. not found / empty)

**`get_data_storage_entry()`**

Get a data storage entry with all details including storage locations and metadata.

```python
get_data_storage_entry(
    data_storage_id: UUID | str,
) -> DataStorageResponse
```

**Returns** `DataStorageResponse`

```json
{
  "data_storage": {
    "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "name": "string",
    "description": "string",
    "content": "string",
    "status": "pending",
    "embedding": [
      0.0
    ],
    "is_collection": false,
    "tags": [
      "string"
    ],
    "parent_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "dataset_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "file_path": "string",
    "bigquery_schema": null,
    "user_id": "string",
    "created_at": "2026-06-03T08:00:00Z",
    "modified_at": "2026-06-03T08:00:00Z",
    "share_status": "private",
    "short_alias": "string"
  },
  "storage_locations": [
    {
      "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "data_storage_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "storage_config": {
        "storage_type": "string",
        "content_type": "string",
        "content_schema": null,
        "metadata": null,
        "location": "string",
        "signed_url": "string"
      },
      "created_at": "2026-06-03T08:00:00Z"
    }
  ]
}
```

**`delete_data_storage_entry()`**

Delete a data storage entry.

```python
delete_data_storage_entry(
    data_storage_entry_id: UUID,
) -> None
```

**Returns** nothing (`None`).

**`search_data_storage()`**

Search data storage objects using structured criteria.

```python
search_data_storage(
    criteria: list[SearchCriterion] | None = None,
    limit: int = 10,
    offset: int = 0,
    filter_logic: FilterLogic = 'OR',
    text_query: str | None = None,
    highlight: bool = False,
    highlight_fields: list[str] | None = None,
    highlight_fragment_size: int = 1000,
    highlight_num_fragments: int = 3,
) -> list[dict]
```

**Arguments**

```json
{
  "criteria": null,
  "limit": 10,
  "offset": 0,
  "filter_logic": "OR",
  "text_query": null,
  "highlight": false,
  "highlight_fields": null,
  "highlight_fragment_size": 1000,
  "highlight_num_fragments": 3
}
```

**Returns** `list[dict]`

*Raw, unvalidated dicts (no Pydantic model).*

**`similarity_search_data_storage()`**

Search data storage objects using vector similarity.

```python
similarity_search_data_storage(
    embedding: list[float],
    size: int = 10,
    min_score: float = 0.7,
    dataset_id: UUID | None = None,
    tags: list[str] | None = None,
    user_id: str | None = None,
    project_id: str | None = None,
) -> list[dict]
```

**Returns** `list[dict]`

*Raw, unvalidated dicts (no Pydantic model).*

**`update_entry_permissions()`**

Update the permissions of a data storage entry.

```python
update_entry_permissions(
    data_storage_id: UUID,
    share_status: ShareStatus,
    permitted_accessors: PermittedAccessors,
) -> DataStorageResponse
```

**Arguments**

```json
{
  "data_storage_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
  "share_status": "private",
  "permitted_accessors": {
    "users": [
      "string"
    ],
    "organizations": [
      "string"
    ]
  }
}
```

**Returns** `DataStorageResponse`

```json
{
  "data_storage": {
    "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "name": "string",
    "description": "string",
    "content": "string",
    "status": "pending",
    "embedding": [
      0.0
    ],
    "is_collection": false,
    "tags": [
      "string"
    ],
    "parent_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "dataset_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "file_path": "string",
    "bigquery_schema": null,
    "user_id": "string",
    "created_at": "2026-06-03T08:00:00Z",
    "modified_at": "2026-06-03T08:00:00Z",
    "share_status": "private",
    "short_alias": "string"
  },
  "storage_locations": [
    {
      "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "data_storage_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "storage_config": {
        "storage_type": "string",
        "content_type": "string",
        "content_schema": null,
        "metadata": null,
        "location": "string",
        "signed_url": "string"
      },
      "created_at": "2026-06-03T08:00:00Z"
    }
  ]
}
```

**`update_entry_tags()`**

Update the tags of a data storage entry.

```python
update_entry_tags(
    data_storage_id: UUID | str,
    tags: list[str],
) -> DataStorageResponse
```

**Returns** `DataStorageResponse`

```json
{
  "data_storage": {
    "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "name": "string",
    "description": "string",
    "content": "string",
    "status": "pending",
    "embedding": [
      0.0
    ],
    "is_collection": false,
    "tags": [
      "string"
    ],
    "parent_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "dataset_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "file_path": "string",
    "bigquery_schema": null,
    "user_id": "string",
    "created_at": "2026-06-03T08:00:00Z",
    "modified_at": "2026-06-03T08:00:00Z",
    "share_status": "private",
    "short_alias": "string"
  },
  "storage_locations": [
    {
      "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "data_storage_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "storage_config": {
        "storage_type": "string",
        "content_type": "string",
        "content_schema": null,
        "metadata": null,
        "location": "string",
        "signed_url": "string"
      },
      "created_at": "2026-06-03T08:00:00Z"
    }
  ]
}
```

#### Async

**`astore_file_content()`**

Awaitable twin of `store_file_content()` — identical arguments and return type.

```python
astore_file_content(
    name: str,
    file_path: str | Path,
    description: str | None = None,
    file_path_override: str | Path | None = None,
    as_collection: bool = False,
    manifest_filename: str | None = None,
    ignore_patterns: list[str] | None = None,
    ignore_filename: str = '.gitignore',
    dataset_id: UUID | None = None,
    project_id: UUID | None = None,
    metadata: dict[str, Any] | None = None,
    tags: list[str] | None = None,
    parent_id: UUID | None = None,
    embedding: list[float] | None = None,
) -> DataStorageResponse
```

**`astore_text_content()`**

Awaitable twin of `store_text_content()` — identical arguments and return type.

```python
astore_text_content(
    name: str,
    content: str,
    description: str | None = None,
    file_path: str | None = None,
    dataset_id: UUID | None = None,
    project_id: UUID | None = None,
    metadata: dict[str, Any] | None = None,
    tags: list[str] | None = None,
    parent_id: UUID | None = None,
    embedding: list[float] | None = None,
) -> DataStorageResponse
```

**`astore_link()`**

Awaitable twin of `store_link()` — identical arguments and return type.

```python
astore_link(
    name: str,
    url: HttpUrl,
    description: str,
    instructions: str,
    api_key: str | None = None,
    metadata: dict[str, Any] | None = None,
    dataset_id: UUID | None = None,
    project_id: UUID | None = None,
    tags: list[str] | None = None,
    parent_id: UUID | None = None,
) -> DataStorageResponse
```

**`afetch_data_from_storage()`**

Awaitable twin of `fetch_data_from_storage()` — identical arguments and return type.

```python
afetch_data_from_storage(
    data_storage_id: UUID | str | None = None,
) -> RawFetchResponse | Path | list[Path] | None
```

**`aget_data_storage_entry()`**

Awaitable twin of `get_data_storage_entry()` — identical arguments and return type.

```python
aget_data_storage_entry(
    data_storage_id: UUID | str,
) -> DataStorageResponse
```

**`adelete_data_storage_entry()`**

Awaitable twin of `delete_data_storage_entry()` — identical arguments and return type.

```python
adelete_data_storage_entry(
    data_storage_entry_id: UUID,
) -> None
```

**`asearch_data_storage()`**

Awaitable twin of `search_data_storage()` — identical arguments and return type.

```python
asearch_data_storage(
    criteria: list[SearchCriterion] | None = None,
    limit: int = 10,
    offset: int = 0,
    filter_logic: FilterLogic = 'OR',
    text_query: str | None = None,
    highlight: bool = False,
    highlight_fields: list[str] | None = None,
    highlight_fragment_size: int = 1000,
    highlight_num_fragments: int = 3,
) -> list[dict]
```

**`asimilarity_search_data_storage()`**

Awaitable twin of `similarity_search_data_storage()` — identical arguments and return type.

```python
asimilarity_search_data_storage(
    embedding: list[float],
    size: int = 10,
    min_score: float = 0.7,
    dataset_id: UUID | None = None,
    tags: list[str] | None = None,
    user_id: str | None = None,
    project_id: str | None = None,
) -> list[dict]
```

**`aupdate_entry_permissions()`**

Awaitable twin of `update_entry_permissions()` — identical arguments and return type.

```python
aupdate_entry_permissions(
    data_storage_id: UUID,
    share_status: ShareStatus,
    permitted_accessors: PermittedAccessors,
) -> DataStorageResponse
```

**`aupdate_entry_tags()`**

Awaitable twin of `update_entry_tags()` — identical arguments and return type.

```python
aupdate_entry_tags(
    data_storage_id: UUID | str,
    tags: list[str],
) -> DataStorageResponse
```

### Projects

#### Sync

**`create_project()`**

Create a new project.

```python
create_project(
    name: str,
    share_status: ShareStatus = 'private',
    permitted_accessors: PermittedAccessors | None = None,
    metadata: dict | None = None,
    description: str | None = None,
    persona_id: UUID | None = None,
) -> UUID
```

**Arguments**

```json
{
  "name": "string",
  "share_status": "private",
  "permitted_accessors": null,
  "metadata": null,
  "description": null,
  "persona_id": null
}
```

**Returns** `UUID`

**`get_project_by_name()`**

Get a project UUID by name.

```python
get_project_by_name(
    name: str,
    limit: int = 2,
    persona_id: UUID | None = None,
) -> UUID | list[UUID]
```

**Returns** `UUID | list[UUID]` — one of:

**`UUID`** — `"019cd179-8d61-7f65-a79c-b965dda9eac3"`

**`list[UUID]`** — `["019cd179-8d61-7f65-a79c-b965dda9eac3"]`

**`add_task_to_project()`**

Add a task to a project. Use this to assign or reassign a task to a project.

```python
add_task_to_project(
    project_id: UUID,
    trajectory_id: str,
) -> None
```

**Returns** nothing (`None`).

**`delete_project()`**

Delete a project.

```python
delete_project(
    project_id: UUID,
    delete_trajectories: bool = True,
) -> None
```

**Returns** nothing (`None`).

#### Async

**`acreate_project()`**

Awaitable twin of `create_project()` — identical arguments and return type.

```python
acreate_project(
    name: str,
    share_status: ShareStatus = 'private',
    permitted_accessors: PermittedAccessors | None = None,
    metadata: dict | None = None,
    description: str | None = None,
    persona_id: UUID | None = None,
) -> UUID
```

**`aget_project_by_name()`**

Awaitable twin of `get_project_by_name()` — identical arguments and return type.

```python
aget_project_by_name(
    name: str,
    limit: int = 2,
    persona_id: UUID | None = None,
) -> UUID | list[UUID]
```

**`aadd_task_to_project()`**

Awaitable twin of `add_task_to_project()` — identical arguments and return type.

```python
aadd_task_to_project(
    project_id: UUID,
    trajectory_id: str,
) -> None
```

**`adelete_project()`**

Awaitable twin of `delete_project()` — identical arguments and return type.

```python
adelete_project(
    project_id: UUID,
    delete_trajectories: bool = True,
) -> None
```


# Edison Client

## Overview

Reference for the objects returned by the most common `EdisonClient` methods.

Every method has a synchronous version and an `a`-prefixed asynchronous twin (e.g. `get_task` / `aget_task`) with identical signatures and return types. Returned objects are Pydantic models unless noted as free-form (raw, unvalidated dicts).

Each page shows the method's arguments and the fields it returns, as JSON.

### Sections

* **Tasks** — submit, fetch, list, cancel, and delete tasks.
* **Data Storage** — upload, fetch, search, and delete stored entries.

##


# run\_tasks\_until\_done()

## `run_tasks_until_done()`

Submit one or more tasks and wait for them to finish

Returns a list of task results, ordered to match the submitted tasks. Each result's shape depends on the job that ran and on the `verbose` flag.

### Arguments

```json
{
  "name": "string",
  "query": "string"
}
```

### Returns

Ordered list of task responses, one per submitted task.

```json
[
  {
    "status": "string",
    "query": "string",
    "user": "string",
    "created_at": "2026-06-03T08:00:00Z",
    "job_name": "string",
    "share_status": "string",
    "permitted_accessors": {},
    "build_owner": "string",
    "environment_name": "string",
    "agent_name": "string",
    "task_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "answer": "string",
    "formatted_answer": "string",
    "answer_reasoning": "string",
    "has_successful_answer": true,
    "total_cost": 0,
    "total_queries": 0
  }
]
```


# create\_task()

## `create_task()`

Submit a task without waiting for it to complete

### Arguments

```json
{
  "name": "string",
  "query": "string"
}
```

### Returns

The trajectory\_id (task ID) of the newly created task.

```json
"019cd179-8d61-7f65-a79c-b965dda9eac3"
```


# get\_task()

## `get_task()`

Fetch the current state of a task

The shape of the result depends on the flags: `lite=true` returns a minimal status payload, `verbose=true` returns the full raw task data, and the default (`verbose=false`) returns the standard task result (with answer fields for jobs that produce one).

### Arguments

```json
{
  "task_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
  "lite": false,
  "verbose": false
}
```

### Returns

Task state; concrete type depends on lite/verbose and job.

> The exact shape varies (`TaskResponse`, `PQATaskResponse`, `TaskResponseVerbose`, `LiteTaskResponse`). The example below shows the most common case.

```json
{
  "status": "string",
  "query": "string",
  "user": "string",
  "created_at": "2026-06-03T08:00:00Z",
  "job_name": "string",
  "share_status": "string",
  "permitted_accessors": {},
  "build_owner": "string",
  "environment_name": "string",
  "agent_name": "string",
  "task_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
  "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
  "answer": "string",
  "formatted_answer": "string",
  "answer_reasoning": "string",
  "has_successful_answer": true,
  "total_cost": 0,
  "total_queries": 0
}
```


# cancel\_task()

## `cancel_task()`

Cancel a task in progress (sync-only, no async twin)

### Arguments

```json
{
  "task_id": "019cd179-8d61-7f65-a79c-b965dda9eac3"
}
```

### Returns

True only if the task was IN\_PROGRESS and is now CANCELLED; False otherwise (e.g. it had already finished).

```json
true
```


# get\_tasks()

## `get_tasks()`

List trajectories with optional filtering

### Arguments

```json
{
  "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
  "name": "string",
  "user": "string",
  "limit": 0,
  "offset": 0,
  "sort_by": "string",
  "sort_order": "string"
}
```

### Returns

Raw, unvalidated dicts (no Pydantic schema). Keys are not guaranteed by a model definition; access defensively.

```json
[
  {
    "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "created_at": "2026-06-03T08:00:00Z",
    "started_at": "2026-06-03T08:00:00Z",
    "crow": "string",
    "user": "string",
    "status": "string",
    "failure_reason": "string",
    "share_status": "string",
    "enabled": true,
    "notification_enabled": true,
    "notification_type": "string",
    "task": "string",
    "task_summary": "string",
    "build_id": "string",
    "gcloud_operation_name": "string",
    "runtime_config": "string",
    "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "project_name": "string",
    "project_description": "string",
    "organization_name": "string",
    "min_estimated_time": 0,
    "max_estimated_time": 0,
    "tags": [
      "string"
    ],
    "pruned_at": "2026-06-03T08:00:00Z"
  }
]
```


# delete\_trajectory()

## `delete_trajectory()`

Delete a task / trajectory

### Arguments

```json
{
  "trajectory_id": "019cd179-8d61-7f65-a79c-b965dda9eac3"
}
```

### Returns

Null on success.

```json
null
```


# store\_file\_content()

## `store_file_content()`

Upload a file's content as a data storage entry

### Arguments

```json
{
  "name": "string",
  "file_path": "string",
  "description": "string"
}
```

### Returns

The stored entry and its backend storage location(s).

```json
{
  "data_storage": {
    "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "name": "string",
    "description": "string",
    "content": "string",
    "status": "string",
    "embedding": [
      0
    ],
    "is_collection": true,
    "tags": [
      "string"
    ],
    "parent_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "dataset_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "file_path": "string",
    "bigquery_schema": null,
    "user_id": "string",
    "created_at": "2026-06-03T08:00:00Z",
    "modified_at": "2026-06-03T08:00:00Z",
    "share_status": "string",
    "short_alias": "string"
  },
  "storage_locations": [
    {
      "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "data_storage_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "storage_config": {
        "storage_type": "string",
        "content_type": "string",
        "content_schema": null,
        "metadata": null,
        "location": "string",
        "signed_url": "string"
      },
      "created_at": "2026-06-03T08:00:00Z"
    }
  ]
}
```


# fetch\_data\_from\_storage()

## `fetch_data_from_storage()`

Download or fetch a stored entry's content

What you get back depends on the storage backend: a GCS-backed entry returns the local filesystem path to the downloaded file (an auto-extracted directory if it was a zip); an entry spread across multiple storage locations returns a list of paths; an inline (raw\_content / pg\_table) entry returns its content directly; and a missing or empty entry returns null. Backend errors or unsupported storage types raise an error.

### Arguments

```json
{
  "data_storage_id": "019cd179-8d61-7f65-a79c-b965dda9eac3"
}
```

### Returns

RawFetchResponse, a Path, a list of Paths, or null.

```json
{
  "filename": "string",
  "content": "string",
  "entry_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
  "entry_name": "string"
}
```


# get\_data\_storage\_entry()

## `get_data_storage_entry()`

Fetch an entry's metadata + storage locations (no download)

### Arguments

```json
{
  "data_storage_id": "019cd179-8d61-7f65-a79c-b965dda9eac3"
}
```

### Returns

Same model as the upload response.

```json
{
  "data_storage": {
    "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "name": "string",
    "description": "string",
    "content": "string",
    "status": "string",
    "embedding": [
      0
    ],
    "is_collection": true,
    "tags": [
      "string"
    ],
    "parent_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "dataset_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "file_path": "string",
    "bigquery_schema": null,
    "user_id": "string",
    "created_at": "2026-06-03T08:00:00Z",
    "modified_at": "2026-06-03T08:00:00Z",
    "share_status": "string",
    "short_alias": "string"
  },
  "storage_locations": [
    {
      "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "data_storage_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
      "storage_config": {
        "storage_type": "string",
        "content_type": "string",
        "content_schema": null,
        "metadata": null,
        "location": "string",
        "signed_url": "string"
      },
      "created_at": "2026-06-03T08:00:00Z"
    }
  ]
}
```


# delete\_data\_storage\_entry()

## `delete_data_storage_entry()`

Delete a data storage entry

### Arguments

```json
{
  "data_storage_entry_id": "019cd179-8d61-7f65-a79c-b965dda9eac3"
}
```

### Returns

Null on success.

```json
null
```


# search\_data\_storage()

## `search_data_storage()`

Search stored entries by criteria and/or a text query

### Arguments

```json
{
  "criteria": {},
  "text_query": "string",
  "limit": 0,
  "offset": 0,
  "filter_logic": "string"
}
```

### Returns

Raw, unvalidated dicts (no Pydantic schema). Keys are not guaranteed by a model definition; access defensively.

```json
[
  {
    "id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "name": "string",
    "description": "string",
    "content": "string",
    "tags": [
      "string"
    ],
    "user_id": "string",
    "dataset_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "parent_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "project_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "project_name": "string",
    "project_description": "string",
    "file_path": "string",
    "is_collection": true,
    "status": "string",
    "share_status": "string",
    "label": "string",
    "path": "string",
    "short_alias": "string",
    "created_at": "2026-06-03T08:00:00Z",
    "modified_at": "2026-06-03T08:00:00Z",
    "metadata": {},
    "embedding": [
      0
    ],
    "score": 0,
    "embedding_similarity": 0,
    "highlights": {},
    "trajectory_id": "019cd179-8d61-7f65-a79c-b965dda9eac3",
    "is_kosmos_project": true
  }
]
```


# Page 1


# docs


# client\_notebook

## Edison platform client usage example

```python
import os

from edison_client import (
    EdisonClient,
    JobNames,
    TaskRequest,
)
from edison_client.models import RuntimeConfig
from ldp.agent import AgentConfig
```

### Client instantiation

Here we instantiate an Edison client and authenticate our access to the platform. By default, the client will use the `EDISON_PLATFORM_API_KEY` environment variable to authenticate. The option `api_key` can be used to pass your API key.

Please log in to the platform and go to your user settings to get your API key.

```python
client = EdisonClient(
    api_key=os.getenv("EDISON_PLATFORM_API_KEY", "your-api-key"),
)
```

### Submit a task to an available Edison job

In the edison platform, we refer to the deployed combination of agent and environment as a `job`. Submitting a task to an edison job is done by calling the `create_task` method, which receives a `TaskRequest` object.

For convenience, one can use the `run_tasks_until_done` method, which submits the task, waits for the task to be completed, and returns a list of `TaskResponse` objects.

```python
task_data = TaskRequest(
    name=JobNames.from_string("literature"),
    query="What is the molecule known to have the greatest solubility in water?",
)
responses = client.run_tasks_until_done(task_data)
task_response = responses[0]

print(f"Job status: {task_response.status}")
print(f"Job answer: \n{task_response.formatted_answer}")
```

You can also pass a `runtime_config` to the `run_tasks_until_done` method, which will be used to configure the agent on runtime. Here, we will define a agent configuration and include it in the `TaskRequest`. This agent is used to decide the next action to take. We will also use the `max_steps` parameter to limit the number of steps the agent will take.

```python
agent = AgentConfig(
    agent_type="SimpleAgent",
    agent_kwargs={
        "model": "gpt-4o",
        "temperature": 0.0,
    },
)
task_data = TaskRequest(
    name=JobNames.LITERATURE,
    query="How many moons does earth have?",
    runtime_config=RuntimeConfig(agent=agent, max_steps=10),
)
responses = client.run_tasks_until_done(task_data)
task_response = responses[0]

print(f"Job status: {task_response.status}")
print(f"Job answer: \n{task_response.formatted_answer}")
```

## Continue a job

The platform allows to ask follow-up questions to the previous job. To accomplish that, we can use the `runtime_config` to pass the `task_id` of the previous task.

Notice that `run_tasks_until_done` accepts both a `TaskRequest` object and a dictionary with keywords arguments.

```python
task_data = TaskRequest(
    name=JobNames.LITERATURE, query="How many species of birds are there?"
)

responses = client.run_tasks_until_done(task_data)
task_response = responses[0]

print(f"First job status: {task_response.status}")
print(f"First job answer: \n{task_response.formatted_answer}")
```

```python
continued_job_data = {
    "name": JobNames.LITERATURE,
    "query": (
        "From the previous answer, specifically, how many species of crows are there?"
    ),
    "runtime_config": {"continued_job_id": task_response.task_id},
}

responses = client.run_tasks_until_done(continued_job_data)
continued_task_response = responses[0]


print(f"Continued job status: {continued_task_response.status}")
print(f"Continued job answer: \n{continued_task_response.formatted_answer}")
```


# Edison Analysis API Tutorial

This notebook provides you with an example usecase for using `Edison Analysis` to perform data analysis.

The only dependency you need to follow along is `edison-client` which you can install via pip:

```bash
pip install edison-client
```

We recommend reading the edison client [docs](https://pypi.org/project/edison-client/) before following this tutorial.

To run a `Edison Analysis` job you should take the following steps:

1. Upload the any artifacts to the data storage service
2. Start an `Edison Analysis` run using the Edison client passing the data storage entry ids along with any other details in the task config
3. Use the output of the task to obtain any data generated by the task

```python
import time

from edison_client import EdisonClient
from edison_client.models import RuntimeConfig, TaskRequest
from edison_client.models.app import JobNames
```

```python
# Instantiate the Edison client with your API key created via the platform
EDISON_API_KEY = ""  # Add your API key here
client = EdisonClient(api_key=EDISON_API_KEY)
```

## File management with Edison Analysis

`Edison Analysis` is designed to run data analysis on files provided by the user or caller. To provide `Edison Analysis` with this data, you'll need to upload it to the Edison data storage service. This service is your one stop shop for sharing, storing and updating data to be used in the Edison ecosystem.

```python
# Uploading a single file to the data storage service
single_file_upload_response = await client.astore_file_content(
    name="Demo file entry for a single file",
    file_path="./datasets/brain_size_data.csv",  # ADD DATASET PATH HERE
    description="This is a test file that will be be analysed by Edison Analysis",
)
```

```python
# Uploading a directory to the data storage service
directory_upload_response = await client.astore_file_content(
    name="Demo file entry for a whole directory",
    file_path="./datasets",  # ADD DATASET FOLDER PATH HERE
    description="This is a directory that will be be analysed by Edison Analysis",
    as_collection=True,
)
```

## Running Your Job

When running a `Edison Analysis` job there are some considerations to take with how you configure the agent. The first things to note are the core configuration settings like `language`, `max_steps` and `query`. In addition to these core settings you have some other options too. The key ones are listed below:

### Additional tools available:

* `query_ensembl`: query the Ensembl database
* `get_convert_gene`: for converting gene IDs from one type to another, for example Ensembl, Entrez, Refseq.
* `search_web`: expose exa.ai (/search) web search as a tool
* `crawl_web`: expose exa.ai (/contents) web crawl as a tool
* `research_web`: expose exa.ai (/research) web research as a tool
* `query_literature`: allow `Edison Analysis` to do calls to `Edison Literature` for literature search
* You can add in either user or system prompt for tool usage. For example: "Use the query\_literature tool to compare your findings against published literature."

### Modifying system prompt

There are two options to modify the system prompt:

1. Replace the existing system prompt completely using `prompting_config["system_prompt"]`
2. Append additional guideline to existing system prompt using `prompting_config["system_prompt_additional_guidelines]`

Build the `prompting_config` dictionary then assign it to the `"prompting_config"` key within `environment_config`

```python
# Define your task
USER_QUERY = "Teach me something new about crows."  # The actual query you want Edison Analysis` to run
SYSTEM_PROMPT = ""  # By setting this, you will replace the system prompt entirely.
SYSTEM_PROMPT_ADDITIONAL_GUIDELINES = (
    "Make all figures in dark mode."  # This will be appended to the system prompt
)
_SYSTEM_PROMPT_CONFIG = {
    "system_prompt": SYSTEM_PROMPT,
    "system_prompt_additional_guidelines": SYSTEM_PROMPT_ADDITIONAL_GUIDELINES,
}
LANGUAGE = "PYTHON"  # Choose between "R" and "PYTHON"
MAX_STEPS = 30  # You can change this to impose a limit on the number of steps the agent can take
```

```python
# Create a task
task_data = TaskRequest(
    name=JobNames.ANALYSIS,
    query=USER_QUERY,
    runtime_config=RuntimeConfig(
        max_steps=MAX_STEPS,
        environment_config={
            "language": LANGUAGE,
            "prompting_config": {
                k: v for k, v in _SYSTEM_PROMPT_CONFIG.items() if v
            },  # See above for documentation
            "data_storage_uris": [
                f"data_entry:{directory_upload_response.data_storage.id}"
            ],
            "additional_tools": None,  # See above for options
        },
    ),
)
trajectory_id = client.create_task(task_data)
print(
    f"Task running on platform, you can view progress live at:https://platform.edisonscientific.com/trajectories/{trajectory_id}"
)
```

```python
# Jobs take on average 3-10 minutes to complete
# We also have inbuilt support for polling, asynchronous tasks and other utilities documented here:
# https://edisonscientific.gitbook.io/edison-cookbook/edison-client
status = "in progress"
while status in {"in progress", "queued"}:
    status = client.get_task(trajectory_id).status
    time.sleep(15)

if status == "failed":
    raise RuntimeError("Task failed")

job_result = client.get_task(trajectory_id, verbose=True)
answer = job_result.environment_frame["state"]["state"]["answer"]
print(f"The agent's answer to your research question is: \n{answer}")
```

## Download Task Output

While the task is executing it will create some artifacts. First the notebook which is where the analysis code will be written and any other artifacts creating during the task.

Once the task has completed you may want to check the contents of the notebook or look through the artifacts generated. To obtain these artifacts, you will need to inspect the output of the agent's final `environment_frame`

```python
output_data = job_result.environment_frame["state"]["info"]["output_data"]
print(output_data)
```

```python
for output_file in output_data:
    download_response = await client.afetch_data_from_storage(
        data_storage_id=output_file["entry_id"]
    )

    # Note there are two potential outcomes here. One where the client downloads
    # the file to your local filesystem if it's above ~10MB. The second is where
    # it will return a RawFetchResponse object which contains the raw content.
    print(download_response)
```


# Best Practices for Optimizing Kosmos Workflows

Kosmos is a data-driven discovery agent. To get the most out of each Kosmos run, we recommend following these guidelines:

## 1. Provide a clear and feasible research objective that requires iteration.

The research objective can be broad and exploratory or focused and hypothesis-driven. While Kosmos can handle multiple objectives, it performs best with a single, well-defined objective. It should have scope for iterative hypothesis generation and testing. The answer should not be obvious after reading a few papers or conducting a single data analysis. A scientific collaborator in a relevant field should be able to reasonably complete the task within weeks or months.‍

## 2. Provide enough context in the research objective.

The research objective should provide sufficient scientific and experimental context as a starting point for data analysis or literature search. While Kosmos is able to conduct literature search and read all provided datasets, it will greatly benefit from context highlighting nuances unique to your research group, such as key experimental design choices, non-obvious assumptions common in the field, or atypical data-handling protocols. Phrase it as you would explain to an experienced colleague who just joined your team.

Here are some examples of ways to improve research objective prompts:

### Inappropriate research objectives

* Which of these tissues have a lower ECM pathway expression score?
* List all the genes that are differentially expressed (p < 0.05) between the 'control' and 'drought' RNA-seq samples.
* How is graphene used for water desalination? Suggest cross-linking agents that are compatible with our method.
* Analyze the attached county-level public health dataset. Does ice cream consumption correlate with asthma rates?

### Good research objectives

* Compare ECM pathway expression levels in the melanoma tissue samples. Propose hypotheses on mechanisms driving these changes and downstream functional consequences.
* Analyze the RNA-seq dataset provided from Arabidopsis thaliana leaves, comparing 'control' (well-watered) to 'drought' (water withholding) samples across a time course. Identify the key regulatory pathways that mediate acclimation. Are there any transcription factor(s) outside the well-known ABA pathway that may be a master regulator of this response?
* We are investigating the use of laminated graphene oxide (GO) membranes for water desalination. A major problem is that these membranes swell and lose selectivity when hydrated. I have attached our recent experimental results showing interlayer spacing vs. salt rejection. Identify the top 3-5 most promising cross-linking agents (e.g., diamines, metal ions) that have been proven to control GO membrane swelling and propose which one would be most compatible with our current layer-by-layer fabrication method.
* The provided county-level public health dataset includes data on disease prevalence, environmental factors, and socio-demographics. Our central hypothesis is that counties with higher air pollution (PM2.5) will have increased asthma prevalence, but we suspect this effect is strongest in low-income communities. Build a regression model to test this hypothesis. Control for confounding variables: this must include population density, average age, and regional differences in pollen, so please treat these as covariates. Propose a follow-up analysis to explore the specific impact of pollen vs. PM2.5.

## 3. Provide sufficient context in the data description.

Ensure that Kosmos has the necessary context about the data provided. For example, if providing a table, ensure all column names are intuitively labeled, or provide an additional sheet that describes what each column name means in more detail. A scientific collaborator in a relevant field should be able to interpret the data with the provided research objective without seeking further clarification.

## 4. Provide complex data.

While Kosmos can operate on simple data (e.g. a csv containing a list of gene names), Kosmos makes the most interesting discoveries when given complex high dimensional data, such as scRNA-seq, proteomics, or environmental parameters across multiple samples or timepoints. You can also provide multiple datasets that are relevant to the research objective. There is no limit on the number of files you provide Kosmos as long as total uncompressed size of your dataset is under 5GB. Kosmos is capable of managing workspaces with 100s of files.

## 5. Provide properly labeled and processed data.

The dataset is used as a starting point for exploratory analysis. While Kosmos can perform quality control steps and is able to correct for artifacts such as batch effects, Kosmos is most powerful when operating on high quality, processed data. We do not recommend getting Kosmos to process your data such as mapping of raw sequencing files or annotating raw imaging data.

## 6. Iterate.

Take some time to write your first Kosmos query so you can avoid the common fallacies listed above. However, trial and error will be your best guide to develop a deep intuition on how to best use the system. Start with a familiar dataset or topic. Since runtime scales with dataset size, begin small for faster results and quicker iteration.


# Best Practices for Interacting with Edison Molecules

Edison Molecules is a chemistry-focused scientific agent. To get the most out of each Edison Molecules run, follow these guidelines for formulating effective queries:

## 1. Be specific about molecular inputs.

When asking about molecules, provide explicit identifiers such as:

* SMILES strings
* CAS numbers
* IUPAC names

Edison Molecules can handle multiple representations, but being explicit reduces ambiguity. A chemist familiar with your field should be able to unambiguously identify the molecule(s) from your query.

| Insufficient queries                         | More detailed queries                                                                                   |
| -------------------------------------------- | ------------------------------------------------------------------------------------------------------- |
| Tell me about anti-inflammatory drugs.       | What are the ADMET properties for aspirin (SMILES: `CC(=O)OC1=CC=CC=C1C(=O)O` or CAS: 50-78-2)?         |
| What are the functional groups of vitamin C? | List the functional groups available in ascorbic acid (SMILES: `C([C@@H]([C@@H]1C(=C(C(=O)O1)O)O)O)O`). |
| Find compounds to be used as painkillers.    | Search for molecules similar to acetaminophen (SMILES: `CC(=O)NC1=CC=C(C=C1)O`) in the ChEMBL database. |

## 2. Specify desired outputs clearly.

Clearly state what you need from Edison Molecules. It can be a synthesis route, molecular property prediction, safety assessment, literature-backed answer, or a combination of these. The more specific you are about the outputs you need, the better Edison Molecules can select appropriate tools and provide actionable results.

| Insufficient queries                          | More detailed queries                                                                                                                                                                |
| --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Tell me about `CN1C=NC2=C1C(=O)N(C(=O)N2C)C`. | What are the ADMET properties for `CN1C=NC2=C1C(=O)N(C(=O)N2C)C`? Also, what are its GHS classification and LD50 value?                                                              |
| Is CC(=O)OC1=CC=CC=C1C(=O)O safe?             | Perform a safety assessment for the molecule with SMILES `CC(=O)OC1=CC=CC=C1C(=O)O`, including GHS classification, LD50 value, chemical weapons screening, and toxicity predictions. |
| Can you help me make aspirin?                 | I need a synthesis route, ADMET property predictions, and a safety assessment for the target molecule with SMILES `CC(=O)OC1=CC=CC=C1C(=O)O`.                                        |

## 3. Use proper chemical terminology and notation.

Employ standard chemical nomenclature and notation. For example, use SMILES representation for reactions: `reactants>reagents>products`. Use standard property names (e.g., ADMET properties) when requesting specific molecular properties. This helps Edison Molecules understand your intent and select the most appropriate computational tools.

| Insufficient queries                           | More detailed queries                                                                                                                                                                                                    |
| ---------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| What happens if I mix ethanol and acetic acid? | Predict the product of the reaction with reaction SMILES: `CCO.CC(=O)O>>CC(=O)OC`. Calculate the reaction enthalpy.                                                                                                      |
| How do I make an ester from alcohol and acid?  | Design a synthesis route for ethyl acetate using reaction SMILES: `CCO.CC(=O)O>>CC(=O)OC` with appropriate catalysts and conditions.                                                                                     |
| Check the drug properties of quercetin.        | Calculate ADMET properties (specifically human intestinal absorption, blood-brain barrier permeability, and cytochrome P450 inhibition) for the molecule with SMILES `C1=CC(=C(C=C1C2=C(C(=O)C3=C(C=C(C=C3O2)O)O)O)O)O`. |

## 4. Break down complex queries.

Structure multi-part questions logically so Edison Molecules can create an effective execution plan. While Edison Molecules can handle multi-step queries and longer workflows, clearly organizing your query helps ensure all components are addressed systematically.

| Insufficient queries                                                  | More detailed queries                                                                                                                                                                                                                                                         |
| --------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Find similar drugs to aspirin.                                        | (1) Analyze aspirin (SMILES: `CC(=O)OC1=CC=CC=C1C(=O)O`) for ADMET properties. (2) Search ChEMBL for similar anti-inflammatory compounds. (3) Perform safety assessments on the top 5 candidates. (4) Propose modifications to improve solubility while maintaining efficacy. |
| I need everything about caffeine and how to make it and what it does. | I need you to give me information about caffeine. Follow these steps: (1) Calculate molecular properties (logP, logS, QED) for SMILES `CN1C=NC2=C1C(=O)N(C(=O)N2C)C`. (2) Design a retrosynthesis route. (3) Predict biological activity targets.                             |
| Find drugs for diabetes, check if they work, and make new ones.       | (1) Search ChEMBL for approved diabetes drugs targeting INSR. (2) Analyze their binding affinities and ADMET properties. (3) Propose 5 novel small molecule candidates with improved properties.                                                                              |

## 5. Provide context when relevant.

Include background information about your use case (e.g., "for drug development" or "for a research synthesis") to help Edison Molecules select appropriate tools and safety considerations. Context about your constraints, goals, or specific requirements enables Edison Molecules to provide more targeted and useful responses.

| Insufficient queries                                  | More detailed queries                                                                                                                                                                                   |
| ----------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Find similar molecules to `CC(=O)OC1=CC=CC=C1C(=O)O`. | Search the ChEMBL database for molecules similar to aspirin (SMILES: `CC(=O)OC1=CC=CC=C1C(=O)O`) for drug repurposing. Return the top 10 candidates with their development phases and bioactivity data. |

## 6. Request specific properties or analyses.

Instead of asking vaguely about a molecule, specify what you need. For example, request specific ADMET properties, synthetic accessibility scores, solubility predictions, or toxicity data. This allows Edison Molecules to use the most appropriate computational tools and provide quantitative, actionable results.

| Insufficient queries                      | More detailed queries                                                                                                                                                                                                       |
| ----------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Modify quercetin to make it more soluble. | Suggest three different substitution modifications to quercetin to make its aqueous solubility (logS) higher. Show me the suggested molecules in your final answer and base your answer in predicted solubility data.       |
| Is aspirin drug-like?                     | Calculate the drug-likeness score (QED), synthetic accessibility score (SAscore), and Lipinski's Rule of Five violations for `CC(=O)OC1=CC=CC=C1C(=O)O`.                                                                    |
| What are the properties of caffeine?      | Calculate the following properties for caffeine (SMILES: `CN1C=NC2=C1C(=O)N(C(=O)N2C)C`): (1) aqueous solubility (logS), (2) partition coefficient (logP), (3) polar surface area (PSA), and (4) number of rotatable bonds. |

## 7. Ask actionable questions that leverage Edison Molecules' toolset.

Frame queries that can be answered using Edison Molecules' computational capabilities rather than pure theoretical discussions without computational support. Edison Molecules excels at property prediction, synthesis planning, reaction analysis, database searches, and literature-enhanced discovery.

| Insufficient queries                              | More detailed queries                                                                                                                                                          |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| How do I make aspirin?                            | Design a retrosynthesis route for the target molecule with SMILES `CC(=O)OC1=CC=CC=C1C(=O)O`. Please identify starting materials and propose reaction steps.                   |
| Can you synthesize caffeine?                      | Propose a synthesis route for caffeine (SMILES: `CN1C=NC2=C1C(=O)N(C(=O)N2C)C`), including retrosynthetic analysis, reactants pricing, and reaction conditions where possible. |
| What drugs treat diabetes?                        | Propose small molecule binders for the insulin receptor (INSR gene symbol). Propose up to 10 candidates and analyze their drug-likeness using QED scores.                      |
| Explain how this ethanol reacts with acetic acid. | Predict the product of the reaction with SMILES `CCO.CC(=O)O>>`, calculate the reaction enthalpy, identify the mechanism, and suggest optimal catalysts and conditions.        |

## 8. Iterate.

Take some time to write your first Edison Molecules query so you can avoid the common pitfalls listed above. Starting with simpler queries will give you faster results and allow quicker iteration, helping you learn how to use Edison Molecules more effectively and get the most out of the system. Your first query does not have to be perfect. Use Edison Molecules interactively to refine your query:

1. Start simple
   * Begin with a well-known molecule or reaction and a small set of properties or a basic retrosynthesis.
2. Inspect the outputs
   * Check if the properties, synthesis steps, or hits match your expectations.
3. Refine your query
   * If you need more detail, add explicit properties, constraints, or additional steps.
   * Example: "Now also calculate logS and propose modifications to improve solubility."
4. Repeat
   * Use the results of one Edison Molecules run as input for the next (e.g., take top hits and ask for safety assessments or optimization ideas).
   * Edison Molecules also accepts follow-up questions to the previous run.


# PaperQA2

[![GitHub](https://img.shields.io/badge/GitHub-black?logo=github\&logoColor=white)](https://github.com/Future-House/paper-qa) [![PyPI version](https://badge.fury.io/py/paper-qa.svg)](https://badge.fury.io/py/paper-qa) [![tests](https://github.com/Future-House/paper-qa/actions/workflows/tests.yml/badge.svg)](https://github.com/Future-House/paper-qa) ![License](https://img.shields.io/badge/License-Apache_2.0-blue.svg) ![PyPI Python Versions](https://img.shields.io/pypi/pyversions/paper-qa)

PaperQA2 is a package for doing high-accuracy retrieval augmented generation (RAG) on PDFs, text files, Microsoft Office documents, and source code files, with a focus on the scientific literature. See our [recent 2024 paper](https://paper.wikicrow.ai) to see examples of PaperQA2's superhuman performance in scientific tasks like question answering, summarization, and contradiction detection.

***

**Table of Contents**

* [Quickstart](#quickstart)
  * [Example Output](#example-output)
* [What is PaperQA2](#what-is-paperqa2)
  * [PaperQA2 vs PaperQA](#paperqa2-vs-paperqa)
  * [PaperQA2 Goes CalVer in December 2025](#paperqa2-goes-calver-in-december-2025)
  * [What's New in Version 5 (aka PaperQA2)?](#whats-new-in-version-5-aka-paperqa2)
  * [What's New in December 2025?](#whats-new-in-december-2025)
  * [PaperQA2 Algorithm](#paperqa2-algorithm)
* [Installation](#installation)
* [CLI Usage](#cli-usage)
  * [Bundled Settings](#bundled-settings)
  * [Rate Limits](#rate-limits)
* [Library Usage](#library-usage)
  * [Agentic Adding/Querying Documents](#agentic-addingquerying-documents)
  * [Manual (No Agent) Adding/Querying Documents](#manual-no-agent-addingquerying-documents)
  * [Async](#async)
  * [Choosing Model](#choosing-model)
    * [Locally Hosted](#locally-hosted)
  * [Embedding Model](#embedding-model)
    * [Specifying the Embedding Model](#specifying-the-embedding-model)
    * [Local Embedding Models (Sentence Transformers)](#local-embedding-models-sentence-transformers)
  * [Adjusting number of sources](#adjusting-number-of-sources)
  * [Using Code or HTML](#using-code-or-html)
  * [Multimodal Support](#multimodal-support)
  * [Using External DB/Vector DB and Caching](#using-external-dbvector-db-and-caching)
  * [Creating Index](#creating-index)
    * [Manifest Files](#manifest-files)
  * [Reusing Index](#reusing-index)
  * [Using Clients Directly](#using-clients-directly)
* [Settings Cheatsheet](#settings-cheatsheet)
* [Where do I get papers?](#where-do-i-get-papers)
* [Callbacks](#callbacks)
  * [Caching Embeddings](#caching-embeddings)
* [Customizing Prompts](#customizing-prompts)
  * [Pre and Post Prompts](#pre-and-post-prompts)
* [FAQ](#faq)
  * [How come I get different results than your papers?](#how-come-i-get-different-results-than-your-papers)
  * [How is this different from LlamaIndex or LangChain?](#how-is-this-different-from-llamaindex-or-langchain)
  * [Can I save or load?](#can-i-save-or-load)
* [Reproduction](#reproduction)
* [Citation](#citation)

***

## Quickstart

In this example we take a folder of research paper PDFs, magically get their metadata - including citation counts with a retraction check, then parse and cache PDFs into a full-text search index, and finally answer the user question with an LLM agent.

```bash
pip install paper-qa
mkdir my_papers
curl -o my_papers/PaperQA2.pdf https://arxiv.org/pdf/2409.13740
cd my_papers
pqa ask 'What is PaperQA2?'
```

### Example Output

Question: Has anyone designed neural networks that compute with proteins or DNA?

> The claim that neural networks have been designed to compute with DNA is supported by multiple sources. The work by Qian, Winfree, and Bruck demonstrates the use of DNA strand displacement cascades to construct neural network components, such as artificial neurons and associative memories, using a DNA-based system (Qian2011Neural pages 1-2, Qian2011Neural pages 15-16, Qian2011Neural pages 54-56). This research includes the implementation of a 3-bit XOR gate and a four-neuron Hopfield associative memory, showcasing the potential of DNA for neural network computation. Additionally, the application of deep learning techniques to genomics, which involves computing with DNA sequences, is well-documented. Studies have applied convolutional neural networks (CNNs) to predict genomic features such as transcription factor binding and DNA accessibility (Eraslan2019Deep pages 4-5, Eraslan2019Deep pages 5-6). These models leverage DNA sequences as input data, effectively using neural networks to compute with DNA. While the provided excerpts do not explicitly mention protein-based neural network computation, they do highlight the use of neural networks in tasks related to protein sequences, such as predicting DNA-protein binding (Zeng2016Convolutional pages 1-2). However, the primary focus remains on DNA-based computation.

## What is PaperQA2

PaperQA2 is engineered to be the best agentic RAG model for working with scientific papers. Here are some features:

* A simple interface to get good answers with grounded responses containing in-text citations.
* State-of-the-art implementation including document metadata-awareness in embeddings and LLM-based re-ranking and contextual summarization (RCS).
* Support for agentic RAG, where a language agent can iteratively refine queries and answers.
* Automatic redundant fetching of paper metadata, including citation and journal quality data from multiple providers.
* A usable full-text search engine for a local repository of PDF/text files.
* A robust interface for customization, with default support for all [LiteLLM](https://docs.litellm.ai/docs/providers) models.

By default, it uses [OpenAI embeddings](https://platform.openai.com/docs/guides/embeddings) and [models](https://platform.openai.com/docs/models) with a Numpy vector DB to embed and search documents. However, you can easily use other closed-source, open-source models or embeddings (see details below).

PaperQA2 depends on some awesome libraries/APIs that make our repo possible. Here are some in no particular order:

1. [Semantic Scholar](https://www.semanticscholar.org/)
2. [Crossref](https://www.crossref.org/)
3. [Unpaywall](https://unpaywall.org/)
4. [Pydantic](https://docs.pydantic.dev/latest/)
5. [tantivy](https://github.com/quickwit-oss/tantivy)
6. [LiteLLM](https://docs.litellm.ai/docs/)
7. [pybtex](https://pybtex.org/)

### PaperQA2 vs PaperQA

We've been working hard on fundamental upgrades for a while and mostly followed [SemVer](https://semver.org/), until [December 2025](#paperqa2-goes-calver-in-december-2025). Meaning we've incremented the major version number on each breaking change. This brings us to the current major version number v5. So why call is the repo now called PaperQA2? We wanted to remark on the fact though that we've exceeded human performance on [many important metrics](https://paper.wikicrow.ai). So we arbitrarily call version 5 and onward PaperQA2, and versions before it as PaperQA1 to denote the significant change in performance. We recognize that we are challenged at naming and counting at FutureHouse, so we reserve the right at any time to arbitrarily change the name to PaperCrow.

### PaperQA2 Goes CalVer in December 2025

Prior to December 2025 we used [semantic versioning](https://semver.org/). This eventually led to confusion in two ways:

1. Developers: should we major version bump based on settings or fundamental system capabilities? What if a bug fix requires breaking changes to the agent's behaviors?
2. Speaking: should one use terminology from our publications (e.g. [PaperQA1](https://arxiv.org/abs/2312.07559), [PaperQA2](https://arxiv.org/abs/2409.13740)) or the Git tags (e.g. v5) from this repo/package? When someone says "PaperQA" -- what version do they mean?

To resolve these confusions, in December 2025, we moved to [calendar versioning](https://calver.org/). The developer burden is diminished because we're basically removing guarantees of backwards compatibility across releases (as CalVer is [ZeroVer](https://0ver.org/) bound to dates). It solves the "speaking" issue because Git tags are now quite different from publication terminology (e.g. PaperQA2 vs `v2025.12.17`). When someone says "PaperQA" it will just refer to the system, not a particular snapshot of agentic behaviors. When someone says "PaperQA2" it will refer to `paper-qa>=5`, which applies to both SemVer tags `v5.0.0` and the new CalVer tags `v2025.12.17`.

This switch is backwards compatible for version 5's SemVer, as the year 2025 is strictly greater than major version 5.

### What's New in Version 5 (aka PaperQA2)?

Version 5 added:

* A CLI `pqa`
* Agentic workflows invoking tools for paper search, gathering evidence, and generating an answer
* Removed much of the statefulness from the `Docs` object
* A migration to LiteLLM for compatibility with many LLM providers as well as centralized rate limits and cost tracking
* A bundled set of configurations (read [this section here](#bundled-settings))) containing known-good hyperparameters

Note that `Docs` objects pickled from prior versions of `PaperQA` are incompatible with version 5, and will need to be rebuilt. Also, our minimum Python version was increased to Python 3.11.

### What's New in December 2025?

The last four months since version `5.29.1` have seen many changes:

* New modalities: tables, figures, non-English languages, math equations
* More and better readers
  * Two new *model-based* PDF readers: [Docling](/paperqa/packages/paper-qa-docling) and [Nvidia nemotron-parse](/paperqa/packages/paper-qa-nemotron)
  * All PDF readers now can parse images and tables, report page numbers, support DPI
  * A reader for Microsoft Office data types
* Multimodal contextual summarization
  * Media objects are also passed to the `summary_llm` during creation
  * Media objects' embedding space is enhanced using an `enrichment_llm` prompt
* Simpler and performant HTTP stack
  * Consolidation from `aiohttp` and `httpx` to just `httpx`
  * Integration with [`httpx-aiohttp`](https://github.com/karpetrosyan/httpx-aiohttp) for performance
* `Context` relevance is simplified and some assumptions were removed
* Many minor features such as retrying `Context` creation upon invalid JSON, compatibility with fall 2025's frontier LLMs, and improved prompt templates
* Multiple fixes in metadata processing via Semantic Scholar and OpenAlex, and metadata processing (e.g. incorrectly inferring identical document IDs for main text and SI)
* Completed the deprecations accrued over the past year

### PaperQA2 Algorithm

To understand PaperQA2, let's start with the pieces of the underlying algorithm. The default workflow of PaperQA2 is as follows:

| Phase                  | PaperQA2 Actions                                                          |
| ---------------------- | ------------------------------------------------------------------------- |
| **1. Paper Search**    | - Get candidate papers from LLM-generated keyword query                   |
|                        | - Chunk, embed, and add candidate papers to state                         |
| **2. Gather Evidence** | - Embed query into vector                                                 |
|                        | - Rank top *k* document chunks in current state                           |
|                        | - Create scored summary of each chunk in the context of the current query |
|                        | - Use LLM to re-score and select most relevant summaries                  |
| **3. Generate Answer** | - Put best summaries into prompt with context                             |
|                        | - Generate answer with prompt                                             |

The tools can be invoked in any order by a language agent. For example, an LLM agent might do a narrow and broad search, or using different phrasing for the gather evidence step from the generate answer step.

## Installation

For a non-development setup, install PaperQA2 (aka version 5) from [PyPI](https://pypi.org/project/paper-qa/). Note version 5 requires Python 3.11+.

```bash
pip install paper-qa>=5
```

For development setup, please refer to the [CONTRIBUTING.md](/paperqa/contributing) file.

PaperQA2 uses an LLM to operate, so you'll need to either set an appropriate [API key environment variable](https://docs.litellm.ai/docs/providers) (i.e. `export OPENAI_API_KEY=sk-...`) or set up an open source LLM server (i.e. using [llamafile](https://github.com/Mozilla-Ocho/llamafile). Any LiteLLM compatible model can be configured to use with PaperQA2.

If you need to index a large set of papers (100+), you will likely want an API key for both [Crossref](https://www.crossref.org/documentation/metadata-plus/metadata-plus-keys/) and [Semantic Scholar](https://www.semanticscholar.org/product/api#api-key), which will allow you to avoid hitting public rate limits using these metadata services. Those can be exported as `CROSSREF_API_KEY` and `SEMANTIC_SCHOLAR_API_KEY` variables.

## CLI Usage

The fastest way to test PaperQA2 is via the CLI. First navigate to a directory with some papers and use the `pqa` cli:

```bash
pqa ask 'What is PaperQA2?'
```

You will see PaperQA2 index your local PDF files, gathering the necessary metadata for each of them (using [Crossref](https://www.crossref.org/) and [Semantic Scholar](https://www.semanticscholar.org/)), search over that index, then break the files into chunked evidence contexts, rank them, and ultimately generate an answer. The next time this directory is queried, your index will already be built (save for any differences detected, like new added papers), so it will skip the indexing and chunking steps.

All prior answers will be indexed and stored, you can view them by querying via the `search` subcommand, or access them yourself in your `PQA_HOME` directory, which defaults to `~/.pqa/`.

```bash
pqa -i 'answers' search 'ranking and contextual summarization'
```

PaperQA2 is highly configurable, when running from the command line, `pqa --help` shows all options and short descriptions. For example to run with a higher temperature:

```bash
pqa --temperature 0.5 ask 'What is PaperQA2?'
```

You can view all settings with `pqa view`. Another useful thing is to change to other templated settings - for example `fast` is a setting that answers more quickly and you can see it with `pqa -s fast view`

Maybe you have some new settings you want to save? You can do that with

```bash
pqa -s my_new_settings --temperature 0.5 --llm foo-bar-5 save
```

and then you can use it with

```bash
pqa -s my_new_settings ask 'What is PaperQA2?'
```

If you run `pqa` with a command which requires a new indexing, say if you change the default chunk\_size, a new index will automatically be created for you.

```bash
pqa --parsing.chunk_size 5000 ask 'What is PaperQA2?'
```

You can also use `pqa` to do full-text search with use of LLMs view the search command. For example, let's save the index from a directory and give it a name:

```bash
pqa -i nanomaterials index
```

Now I can search for papers about thermoelectrics:

```bash
pqa -i nanomaterials search thermoelectrics
```

or I can use the normal ask

```bash
pqa -i nanomaterials ask 'Are there nm scale features in thermoelectric materials?'
```

Both the CLI and module have pre-configured settings based on prior performance and our publications, they can be invoked as follows:

```bash
pqa --settings <setting name> \
    ask 'Are there nm scale features in thermoelectric materials?'
```

### Bundled Settings

Inside [`src/paperqa/configs`](https://github.com/Future-House/paper-qa/blob/main/src/paperqa/configs/README.md) we bundle known useful settings:

| Setting Name  | Description                                                                                                                  |
| ------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| high\_quality | Highly performant, relatively expensive (due to having `evidence_k` = 15) query using a `ToolSelector` agent.                |
| fast          | Setting to get answers cheaply and quickly.                                                                                  |
| wikicrow      | Setting to emulate the Wikipedia article writing used in our WikiCrow publication.                                           |
| contracrow    | Setting to find contradictions in papers, your query should be a claim that needs to be flagged as a contradiction (or not). |
| debug         | Setting useful solely for debugging, but not in any actual application beyond debugging.                                     |
| tier1\_limits | Settings that match OpenAI rate limits for each tier, you can use `tier<1-5>_limits` to specify the tier.                    |

### Rate Limits

If you are hitting rate limits, say with the OpenAI Tier 1 plan, you can add them into PaperQA2. For each OpenAI tier, a pre-built setting exists to limit usage.

```bash
pqa --settings 'tier1_limits' ask 'What is PaperQA2?'
```

This will limit your system to use the [tier1\_limits](https://github.com/Future-House/paper-qa/blob/main/src/paperqa/configs/tier1_limits.json), and slow down your queries to accommodate.

You can also specify them manually with any rate limit string that matches the specification in the [limits](https://limits.readthedocs.io/en/stable/quickstart.html#rate-limit-string-notation) module:

```bash
pqa --summary_llm_config '{"rate_limit": {"gpt-4o-2024-11-20": "30000 per 1 minute"}}' \
    ask 'What is PaperQA2?'
```

Or by adding into a `Settings` object, if calling imperatively:

```python
from paperqa import Settings, ask

answer_response = ask(
    "What is PaperQA2?",
    settings=Settings(
        llm_config={"rate_limit": {"gpt-4o-2024-11-20": "30000 per 1 minute"}},
        summary_llm_config={"rate_limit": {"gpt-4o-2024-11-20": "30000 per 1 minute"}},
    ),
)
```

## Library Usage

PaperQA2's full workflow can be accessed via Python directly:

```python
from paperqa import Settings, ask

answer_response = ask(
    "What is PaperQA2?",
    settings=Settings(temperature=0.5, paper_directory="my_papers"),
)
```

Please see our [installation docs](#installation) for how to install the package from PyPI.

### Agentic Adding/Querying Documents

The answer object has the following attributes: `formatted_answer`, `answer` (answer alone), `question` , and `context` (the summaries of passages found for answer). `ask` will use the `SearchPapers` tool, which will query a local index of files, you can specify this location via the `Settings` object:

```python
from paperqa import Settings, ask

answer_response = ask(
    "What is PaperQA2?",
    settings=Settings(
        temperature=0.5, agent={"index": {"paper_directory": "my_papers"}}
    ),
)
```

`ask` is just a convenience wrapper around the real entrypoint, which can be accessed if you'd like to run concurrent asynchronous workloads:

```python
from paperqa import Settings, agent_query

answer_response = await agent_query(
    query="What is PaperQA2?",
    settings=Settings(
        temperature=0.5, agent={"index": {"paper_directory": "my_papers"}}
    ),
)
```

The default agent will use an LLM based agent, but you can also specify a `"fake"` agent to use a hard coded call path of search -> gather evidence -> answer to reduce token usage.

### Manual (No Agent) Adding/Querying Documents

Normally via agent execution, the agent invokes the search tool, which adds documents to the `Docs` object for you behind the scenes. However, if you prefer fine-grained control, you can directly interact with the `Docs` object.

Note that manually adding and querying `Docs` does not impact performance. It just removes the automation associated with an agent picking the documents to add.

```python
from paperqa import Docs, Settings

# valid extensions include .pdf, .txt, .md, .html, .docx, .xlsx, .pptx, and code files (e.g., .py, .ts, .yaml)
doc_paths = ("myfile.pdf", "myotherfile.pdf")

# Prepare the Docs object by adding a bunch of documents
docs = Docs()
for doc_path in doc_paths:
    await docs.aadd(doc_path)

# Set up how we want to query the Docs object
settings = Settings()
settings.llm = "claude-3-5-sonnet-20240620"
settings.answer.answer_max_sources = 3

# Query the Docs object to get an answer
session = await docs.aquery("What is PaperQA2?", settings=settings)
print(session)
```

### Async

PaperQA2 is written to be used asynchronously. The synchronous API is just a wrapper around the async. Here are the methods and their `async` equivalents:

| Sync                | Async                |
| ------------------- | -------------------- |
| `Docs.add`          | `Docs.aadd`          |
| `Docs.add_file`     | `Docs.aadd_file`     |
| `Docs.add_url`      | `Docs.aadd_url`      |
| `Docs.get_evidence` | `Docs.aget_evidence` |
| `Docs.query`        | `Docs.aquery`        |

The synchronous version just calls the async version in a loop. Most modern python environments support `async` natively (including Jupyter notebooks!). So you can do this in a Jupyter Notebook:

```python
import asyncio
from paperqa import Docs


async def main() -> None:
    docs = Docs()
    # valid extensions include .pdf, .txt, .md, .html, .docx, .xlsx, .pptx, and code files (e.g., .py, .ts, .yaml)
    for doc in ("myfile.pdf", "myotherfile.pdf"):
        await docs.aadd(doc)

    session = await docs.aquery("What is PaperQA2?")
    print(session)


asyncio.run(main())
```

### Choosing Model

By default, PaperQA2 uses OpenAI's `gpt-4o-2024-11-20` model for the `summary_llm`, `llm`, and `agent_llm`. Please see the [Settings Cheatsheet](#settings-cheatsheet) for more information on these settings. PaperQA2 also defaults to using OpenAI's `text-embedding-3-small` model for the `embedding` setting. If you don't have an OpenAI API key, you can use a different embedding model. More information about embedding models can be found [in the "Embedding Model" section](#embedding-model).

We use the [`lmi`](https://github.com/Future-House/ldp/tree/main/packages/lmi) package for our LLM interface, which in turn uses `litellm` to support many LLM providers. You can adjust this easily to use any model supported by `litellm`:

```python
from paperqa import Settings, ask

answer_response = ask(
    "What is PaperQA2?",
    settings=Settings(
        llm="gpt-4o-mini", summary_llm="gpt-4o-mini", agent={"index": {"paper_directory": "my_papers"}}
    ),
)
```

To use Claude, make sure you set the `ANTHROPIC_API_KEY` environment variable. In this example, we also use a different embedding model. Please make sure to `pip install paper-qa[local]` to use a local embedding model.

```python
from paperqa import Settings, ask
from paperqa.settings import AgentSettings

answer_response = ask(
    "What is PaperQA2?",
    settings=Settings(
        llm="claude-3-5-sonnet-20240620",
        summary_llm="claude-3-5-sonnet-20240620",
        agent=AgentSettings(agent_llm="claude-3-5-sonnet-20240620"),
        # SEE: https://huggingface.co/sentence-transformers/multi-qa-MiniLM-L6-cos-v1
        embedding="st-multi-qa-MiniLM-L6-cos-v1",
    ),
)
```

Or Gemini, by setting the `GEMINI_API_KEY` from Google AI Studio

```python
from paperqa import Settings, ask
from paperqa.settings import AgentSettings

answer_response = ask(
    "What is PaperQA2?",
    settings=Settings(
        llm="gemini/gemini-2.0-flash",
        summary_llm="gemini/gemini-2.0-flash",
        agent=AgentSettings(agent_llm="gemini/gemini-2.0-flash"),
        embedding="gemini/text-embedding-004",
    ),
)
```

#### Locally Hosted

You can use llama.cpp to be the LLM. Note that you should be using relatively large models, because PaperQA2 requires following a lot of instructions. You won't get good performance with 7B models.

The easiest way to get set-up is to download a [llama file](https://github.com/Mozilla-Ocho/llamafile) and execute it with `-cb -np 4 -a my-llm-model --embedding` which will enable continuous batching and embeddings.

```python
from paperqa import Settings, ask

local_llm_config = dict(
    model_list=[
        dict(
            model_name="my_llm_model",
            litellm_params=dict(
                model="my-llm-model",
                api_base="http://localhost:8080/v1",
                api_key="sk-no-key-required",
                temperature=0.1,
                frequency_penalty=1.5,
                max_tokens=512,
            ),
        )
    ]
)

answer_response = ask(
    "What is PaperQA2?",
    settings=Settings(
        llm="my-llm-model",
        llm_config=local_llm_config,
        summary_llm="my-llm-model",
        summary_llm_config=local_llm_config,
    ),
)
```

Models hosted with `ollama` are also supported. To run the example below make sure you have downloaded llama3.2 and mxbai-embed-large via ollama.

```python
from paperqa import Settings, ask

local_llm_config = {
    "model_list": [
        {
            "model_name": "ollama/llama3.2",
            "litellm_params": {
                "model": "ollama/llama3.2",
                "api_base": "http://localhost:11434",
            },
        }
    ]
}

answer_response = ask(
    "What is PaperQA2?",
    settings=Settings(
        llm="ollama/llama3.2",
        llm_config=local_llm_config,
        summary_llm="ollama/llama3.2",
        summary_llm_config=local_llm_config,
        embedding="ollama/mxbai-embed-large",
    ),
)
```

### Embedding Model

Embeddings are used to retrieve k texts (where k is specified via `Settings.answer.evidence_k`) for re-ranking and contextual summarization. If you don't want to use embeddings, but instead just fetch all chunks, disable "evidence retrieval" via the `Settings.answer.evidence_retrieval` setting.

PaperQA2 defaults to using OpenAI (`text-embedding-3-small`) embeddings, but has flexible options for both vector stores and embedding choices.

#### Specifying the Embedding Model

The simplest way to specify the embedding model is via `Settings.embedding`:

```python
from paperqa import Settings, ask

answer_response = ask(
    "What is PaperQA2?",
    settings=Settings(embedding="text-embedding-3-large"),
)
```

`embedding` accepts any embedding model name supported by litellm. PaperQA2 also supports an embedding input of `"hybrid-<model_name>"` i.e. `"hybrid-text-embedding-3-small"` to use a hybrid sparse keyword (based on a token modulo embedding) and dense vector embedding, where any litellm model can be used in the dense model name. `"sparse"` can be used to use a sparse keyword embedding only.

Embedding models are used to create PaperQA2's index of the full-text embedding vectors (`texts_index` argument). The embedding model can be specified as a setting when you are adding new papers to the `Docs` object:

```python
from paperqa import Docs, Settings

docs = Docs()
for doc in ("myfile.pdf", "myotherfile.pdf"):
    await docs.aadd(doc, settings=Settings(embedding="text-embedding-large-3"))
```

Note that PaperQA2 uses Numpy as a dense vector store. Its design of using a keyword search initially reduces the number of chunks needed for each answer to a relatively small number < 1k. Therefore, `NumpyVectorStore` is a good place to start, it's a simple in-memory store, without an index. However, if a larger-than-memory vector store is needed, you can an external vector database like [Qdrant](https://qdrant.tech/) via the `QdrantVectorStore` class.

The hybrid embeddings can be customized:

```python
from paperqa import (
    Docs,
    HybridEmbeddingModel,
    SparseEmbeddingModel,
    LiteLLMEmbeddingModel,
)


model = HybridEmbeddingModel(
    models=[LiteLLMEmbeddingModel(), SparseEmbeddingModel(ndim=1024)]
)
docs = Docs()
for doc in ("myfile.pdf", "myotherfile.pdf"):
    await docs.aadd(doc, embedding_model=model)
```

The sparse embedding (keyword) models default to having 256 dimensions, but this can be specified via the `ndim` argument.

#### Local Embedding Models (Sentence Transformers)

You can use a `SentenceTransformerEmbeddingModel` model if you install `sentence-transformers`, which is [a local embedding library](https://sbert.net/) with support for HuggingFace models and more. You can install it by adding the `local` extras.

```sh
pip install paper-qa[local]
```

and then prefix embedding model names with `st-`:

```python
from paperqa import Settings, ask

answer_response = ask(
    "What is PaperQA2?",
    settings=Settings(embedding="st-multi-qa-MiniLM-L6-cos-v1"),
)
```

or with a hybrid model

```python
from paperqa import Settings, ask

answer_response = ask(
    "What is PaperQA2?",
    settings=Settings(embedding="hybrid-st-multi-qa-MiniLM-L6-cos-v1"),
)
```

### Adjusting number of sources

You can adjust the numbers of sources (passages of text) to reduce token usage or add more context. `k` refers to the top k most relevant and diverse (may from different sources) passages. Each passage is sent to the LLM to summarize, or determine if it is irrelevant. After this step, a limit of `max_sources` is applied so that the final answer can fit into the LLM context window. Thus, `k` > `max_sources` and `max_sources` is the number of sources used in the final answer.

```python
from paperqa import Settings

settings = Settings()
settings.answer.answer_max_sources = 3
settings.answer.evidence_k = 5

await docs.aquery(
    "What is PaperQA2?",
    settings=settings,
)
```

### Using Code or HTML

You do not need to use papers -- you can use code or raw HTML. Note that this tool is focused on answering questions, so it won't do well at writing code. One note is that the tool cannot infer citations from code, so you will need to provide them yourself.

```python
import glob
import os
from paperqa import Docs

source_files = glob.glob("**/*.js")

docs = Docs()
for f in source_files:
    # this assumes the file names are unique in code
    await docs.aadd(
        f, citation="File " + os.path.basename(f), docname=os.path.basename(f)
    )
session = await docs.aquery("Where is the search bar in the header defined?")
print(session)
```

### Multimodal Support

Multimodal support centers on:

* Standalone images
* Images or tables in PDFs

The `Docs` object stores media via a `ParsedMedia` object. When chunking a document, media are not split at chunk boundaries, so it's possible 2+ chunks can correspond with the same media. This means within PaperQA each chunk has a one-to-many relationship between `ParsedMedia` and chunks.

Depending on the source document, the same image can appear multiple times (e.g. each page of a PDF has a logo in the margins). Thus, clients should consider media databases to have a many-to-many relationship with chunks.

Since PaperQA's evidence gathering process centers on text-based retrieval, it's possible relevant image(s) or table(s) aren't retrieved because their associated text content is irrelevant. For a concrete example, imagine the figure in a paper has a terse caption and is placed one page after relevant main-text discussion. To solve this problem, PaperQA supports media enrichment at document read-time. Basically after reading in the PDF, the `parsing.enrichment_llm` is given the `parsing.enrichment_prompt` and co-located text to generate a synthetic caption for every image/table. The synthetic captions are used to shift the embeddings of each text chunk, but are kept separate from the actual source text. This way evidence gathering can fetch relevant images/tables without risk of polluting contextual summaries with LLM-generated captions.

If you want multimodal PDF reading, but do not want enrichment (since adds one LLM prompt/media at read-time), enrichment can be disabled by setting `parsing.multimodal` to `ON_WITHOUT_ENRICHMENT`.

When creating contextual summaries on a given chunk (a `Text`), the summary LLM is passed both the chunk's text and the chunk's associated media, but the output contextual summary itself remains text-only.

If you would like, specifying the prompt `paperqa.prompts.summary_json_multimodal_system_prompt` to the setting `prompt.summary_json_system` will include a `used_images` flag attributing usage of images in any contextual summarizations.

### Using External DB/Vector DB and Caching

You may want to cache parsed texts and embeddings in an external database or file. You can then build a Docs object from those directly:

```python
from paperqa import Docs, Doc, Text

docs = Docs()

for ... in my_docs:
    doc = Doc(docname=..., citation=..., dockey=..., citation=...)
    texts = [Text(text=..., name=..., doc=doc) for ... in my_texts]
    docs.add_texts(texts, doc)
```

### Creating Index

Indexes will be placed in the [home directory](https://docs.python.org/3/library/pathlib.html#pathlib.Path.home) by default. This can be controlled via the `PQA_HOME` environment variable.

Indexes are made by reading files in the `IndexSettings.paper_directory`. By default, we recursively read from subdirectories of the paper directory, unless disabled using `IndexSettings.recurse_subdirectories`. The paper directory is not modified in any way, it's just read from.

#### Manifest Files

The indexing process attempts to infer paper metadata like title and DOI using LLM-powered text processing. You can avoid this point of uncertainty using a "manifest" file, which is a CSV containing `DocDetails` fields (order doesn't matter). For example:

* `file_location`: relative path to the paper's PDF within the index directory
* `doi`: DOI of the paper
* `title`: title of the paper

By providing this information, we ensure queries to metadata providers like Crossref are accurate.

To ease creating a manifest, there is a helper class method `Doc.to_csv`, which also works when called on `DocDetails`.

### Reusing Index

The local search indexes are built based on a hash of the current `Settings` object. So make sure you properly specify the `paper_directory` to your `IndexSettings` object. In general, it's advisable to:

1. Pre-build an index given a folder of papers (can take several minutes)
2. Reuse the index to perform many queries

```python
import os

from paperqa import Settings
from paperqa.agents.main import agent_query
from paperqa.agents.search import get_directory_index


async def amain(folder_of_papers: str | os.PathLike) -> None:
    settings = Settings(agent={"index": {"paper_directory": folder_of_papers}})

    # 1. Build the index. Note an index name is autogenerated when unspecified
    built_index = await get_directory_index(settings=settings)
    print(settings.get_index_name())  # Display the autogenerated index name
    print(await built_index.index_files)  # Display the index contents

    # 2. Use the settings as many times as you want with ask
    answer_response_1 = await agent_query(
        query="What is a cool retrieval augmented generation technique?",
        settings=settings,
    )
    answer_response_2 = await agent_query(
        query="What is PaperQA2?",
        settings=settings,
    )
```

### Using Clients Directly

One of the most powerful features of PaperQA2 is its ability to combine data from multiple metadata sources. For example, [Unpaywall](https://unpaywall.org/) can provide open access status/direct links to PDFs, [Crossref](https://www.crossref.org/) can provide bibtex, and [Semantic Scholar](https://www.semanticscholar.org/) can provide citation licenses. Here's a short demo of how to do this:

```python
from paperqa.clients import DocMetadataClient, ALL_CLIENTS

client = DocMetadataClient(metadata_clients=ALL_CLIENTS)
details = await client.query(title="Augmenting language models with chemistry tools")

print(details.formatted_citation)
# Andres M. Bran, Sam Cox, Oliver Schilter, Carlo Baldassari,
# Andrew D. White, and Philippe Schwaller.
#  Augmenting large language models with chemistry tools. Nature Machine Intelligence,
# 6:525-535, May 2024. URL: https://doi.org/10.1038/s42256-024-00832-8,
# doi:10.1038/s42256-024-00832-8.
# This article has 243 citations and is from a domain leading peer-reviewed journal.

print(details.citation_count)
# 243

print(details.license)
# cc-by

print(details.pdf_url)
# https://www.nature.com/articles/s42256-024-00832-8.pdf
```

the `client.query` is meant to check for exact matches of title. It's a bit robust (like to casing, missing a word). There are duplicates for titles though - so you can also add authors to disambiguate. Or you can provide a doi directly `client.query(doi="10.1038/s42256-024-00832-8")`.

If you're doing this at a large scale, you may not want to use `ALL_CLIENTS` (just omit the argument) and you can specify which specific fields you want to speed up queries. For example:

```python
details = await client.query(
    title="Augmenting large language models with chemistry tools",
    authors=["Andres M. Bran", "Sam Cox"],
    fields=["title", "doi"],
)
```

will return much faster than the first query and we'll be certain the authors match.

## Settings Cheatsheet

| Setting                                      | Default                                | Description                                                                                                                    |
| -------------------------------------------- | -------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| `llm`                                        | `"gpt-4o-2024-11-20"`                  | LLM for general use including metadata inference (see Docs.aadd) and answer generation (see Docs.aquery and gen\_answer tool). |
| `llm_config`                                 | `None`                                 | Optional configuration for `llm`.                                                                                              |
| `summary_llm`                                | `"gpt-4o-2024-11-20"`                  | LLM for creating contextual summaries (see Docs.aget\_evidence and gather\_evidence tool).                                     |
| `summary_llm_config`                         | `None`                                 | Optional configuration for `summary_llm`.                                                                                      |
| `embedding`                                  | `"text-embedding-3-small"`             | Embedding model for embedding text chunks when adding papers.                                                                  |
| `embedding_config`                           | `None`                                 | Optional configuration for `embedding`.                                                                                        |
| `temperature`                                | `0.0`                                  | Temperature for LLMs.                                                                                                          |
| `batch_size`                                 | `1`                                    | Batch size for calling LLMs.                                                                                                   |
| `texts_index_mmr_lambda`                     | `1.0`                                  | Lambda for MMR in text index.                                                                                                  |
| `verbosity`                                  | `0`                                    | Integer verbosity level for logging (0-3). 3 = all LLM/Embeddings calls logged.                                                |
| `custom_context_serializer`                  | `None`                                 | Custom async function (see typing for signature) to override the default answer context serialization.                         |
| `answer.evidence_k`                          | `10`                                   | Number of evidence pieces to retrieve.                                                                                         |
| `answer.evidence_retrieval`                  | `True`                                 | Use retrieval vs processing all docs.                                                                                          |
| `answer.evidence_summary_length`             | `"about 100 words"`                    | Length of evidence summary.                                                                                                    |
| `answer.evidence_skip_summary`               | `False`                                | Whether to skip summarization.                                                                                                 |
| `answer.evidence_text_only_fallback`         | `False`                                | Whether to allow context creation to retry without media present.                                                              |
| `answer.answer_max_sources`                  | `5`                                    | Max number of sources for an answer.                                                                                           |
| `answer.max_answer_attempts`                 | `None`                                 | Max attempts to generate an answer.                                                                                            |
| `answer.answer_length`                       | `"about 200 words, but can be longer"` | Length of final answer.                                                                                                        |
| `answer.max_concurrent_requests`             | `4`                                    | Max concurrent requests to LLMs.                                                                                               |
| `answer.answer_filter_extra_background`      | `False`                                | Whether to cite background info from model.                                                                                    |
| `answer.get_evidence_if_no_contexts`         | `True`                                 | Allow lazy evidence gathering.                                                                                                 |
| `answer.group_contexts_by_question`          | `False`                                | Groups the final contexts by the underlying `gather_evidence` question in the final context prompt.                            |
| `answer.evidence_relevance_score_cutoff`     | `1`                                    | Cutoff evidence relevance score to include in the answer context (inclusive)                                                   |
| `answer.skip_evidence_citation_strip`        | `False`                                | Skip removal of citations from the `gather_evidence` contexts                                                                  |
| `parsing.page_size_limit`                    | `1,280,000`                            | Character limit per page.                                                                                                      |
| `parsing.use_doc_details`                    | `True`                                 | Whether to get metadata details for docs.                                                                                      |
| `parsing.reader_config`                      | `dict`                                 | Optional keyword arguments for the document reader.                                                                            |
| `parsing.multimodal`                         | `True`                                 | Control to parse both text and media from applicable documents, as well as potentially enriching them with text descriptions.  |
| `parsing.defer_embedding`                    | `False`                                | Whether to defer embedding until summarization.                                                                                |
| `parsing.parse_pdf`                          | `paperqa_pypdf.parse_pdf_to_pages`     | Function to parse PDF files.                                                                                                   |
| `parsing.configure_pdf_parser`               | No-op                                  | Callable to configure the PDF parser within `parse_pdf`, useful for behaviors such as enabling logging.                        |
| `parsing.doc_filters`                        | `None`                                 | Optional filters for allowed documents.                                                                                        |
| `parsing.use_human_readable_clinical_trials` | `False`                                | Parse clinical trial JSONs into readable text.                                                                                 |
| `parsing.enrichment_llm`                     | `"gpt-4o-2024-11-20"`                  | LLM for media enrichment.                                                                                                      |
| `parsing.enrichment_llm_config`              | `None`                                 | Optional configuration for `enrichment_llm`.                                                                                   |
| `parsing.enrichment_page_radius`             | `1`                                    | Page radius for context text in enrichment.                                                                                    |
| `parsing.enrichment_prompt`                  | `image_enrichment_prompt_template`     | Prompt template for enriching media.                                                                                           |
| `parsing.citation_prompt`                    | `citation_prompt`                      | Prompt to create citation from peeking one chunk.                                                                              |
| `parsing.structured_citation_prompt`         | `structured_citation_prompt`           | Prompt to create a citation (in JSON) from peeking one chunk.                                                                  |
| `parsing.disable_doc_valid_check`            | `False`                                | Flag to disable checking if a document looks like text (was parsed correctly).                                                 |
| `prompts.summary`                            | `summary_prompt`                       | Template for summarizing text, must contain variables matching `summary_prompt`.                                               |
| `prompts.qa`                                 | `qa_prompt`                            | Template for QA, must contain variables matching `qa_prompt`.                                                                  |
| `prompts.select`                             | `select_paper_prompt`                  | Template for selecting papers, must contain variables matching `select_paper_prompt`.                                          |
| `prompts.pre`                                | `None`                                 | Optional pre-prompt templated with just the original question to append information before a qa prompt.                        |
| `prompts.post`                               | `None`                                 | Optional post-processing prompt that can access PQASession fields.                                                             |
| `prompts.system`                             | `default_system_prompt`                | System prompt for the model.                                                                                                   |
| `prompts.use_json`                           | `True`                                 | Whether to use JSON formatting.                                                                                                |
| `prompts.summary_json`                       | `summary_json_prompt`                  | JSON-specific summary prompt.                                                                                                  |
| `prompts.summary_json_system`                | `summary_json_system_prompt`           | System prompt for JSON summaries.                                                                                              |
| `prompts.context_outer`                      | `CONTEXT_OUTER_PROMPT`                 | Prompt for how to format all contexts in generate answer.                                                                      |
| `prompts.context_inner`                      | `CONTEXT_INNER_PROMPT`                 | Prompt for how to format a single context in generate answer. Must contain 'name' and 'text' variables.                        |
| `prompts.answer_iteration_prompt`            | `answer_iteration_prompt_template`     | Prompt to inject existing prior answers to allow iteration. Default injects no prior answers.                                  |
| `agent.agent_llm`                            | `"gpt-4o-2024-11-20"`                  | LLM inside the agent making tool selections.                                                                                   |
| `agent.agent_llm_config`                     | `None`                                 | Optional configuration for `agent_llm`.                                                                                        |
| `agent.agent_type`                           | `"ToolSelector"`                       | Type of agent to use.                                                                                                          |
| `agent.agent_config`                         | `None`                                 | Optional kwarg for AGENT constructor.                                                                                          |
| `agent.agent_system_prompt`                  | `env_system_prompt`                    | Optional system prompt message.                                                                                                |
| `agent.agent_prompt`                         | `env_reset_prompt`                     | Agent prompt.                                                                                                                  |
| `agent.return_paper_metadata`                | `False`                                | Whether to include paper title/year in search tool results.                                                                    |
| `agent.search_count`                         | `8`                                    | Search count.                                                                                                                  |
| `agent.timeout`                              | `500.0`                                | Timeout on agent execution (seconds).                                                                                          |
| `agent.tool_names`                           | `None`                                 | Optional override on tools to provide the agent.                                                                               |
| `agent.max_timesteps`                        | `None`                                 | Optional upper limit on environment steps.                                                                                     |
| `agent.agent_evidence_n`                     | `1`                                    | Top n ranked evidences shown to the agent after gathering evidence.                                                            |
| `agent.rebuild_index`                        | `True`                                 | Flag to rebuild the index at the start of agent runners.                                                                       |
| `agent.callbacks`                            | `{}`                                   | Named lists of callables to be invoked with environment state.                                                                 |
| `agent.index.name`                           | `None`                                 | Optional name of the index.                                                                                                    |
| `agent.index.paper_directory`                | `Current working directory`            | Directory containing papers to be indexed.                                                                                     |
| `agent.index.manifest_file`                  | `None`                                 | Path to manifest CSV with document attributes.                                                                                 |
| `agent.index.index_directory`                | `pqa_directory("indexes")`             | Directory to store PQA indexes.                                                                                                |
| `agent.index.use_absolute_paper_directory`   | `False`                                | Whether to use absolute paper directory path.                                                                                  |
| `agent.index.recurse_subdirectories`         | `True`                                 | Whether to recurse into subdirectories when indexing.                                                                          |
| `agent.index.concurrency`                    | `5`                                    | Number of concurrent filesystem reads.                                                                                         |
| `agent.index.sync_with_paper_directory`      | `True`                                 | Whether to sync index with paper directory on load.                                                                            |
| `agent.index.batch_size`                     | `1`                                    | Number of files to process before committing to the index.                                                                     |
| `agent.index.files_filter`                   | `lambda f: f.suffix in {...}`          | Filter function to mark files in the paper directory to index.                                                                 |

## Where do I get papers?

Well that's a really good question! It's probably best to just download PDFs of papers you think will help answer your question and start from there.

See detailed docs [about zotero, openreview and parsing](/paperqa/docs/tutorials/where_do_i_get_papers)

## Callbacks

To execute a function on each chunk of LLM completions, you need to provide a function that can be executed on each chunk. For example, to get a typewriter view of the completions, you can do:

```python
from paperqa import Docs


def typewriter(chunk: str) -> None:
    print(chunk, end="")


docs = Docs()

# add some docs...

await docs.aquery("What is PaperQA2?", callbacks=[typewriter])
```

### Caching Embeddings

In general, embeddings are cached when you pickle a `Docs` regardless of what vector store you use. So as long as you save your underlying `Docs` object, you should be able to avoid re-embedding your documents.

## Customizing Prompts

You can customize any of the prompts using settings.

```python
from paperqa import Docs, Settings

my_qa_prompt = (
    "Answer the question '{question}'\n"
    "Use the context below if helpful. "
    "You can cite the context using the key like (pqac-abcd1234). "
    "If there is insufficient context, write a poem "
    "about how you cannot answer.\n\n"
    "Context: {context}"
)

docs = Docs()
settings = Settings()
settings.prompts.qa = my_qa_prompt
await docs.aquery("What is PaperQA2?", settings=settings)
```

### Pre and Post Prompts

Following the syntax above, you can also include prompts that are executed after the query and before the query. For example, you can use this to critique the answer.

## FAQ

### How come I get different results than your papers?

Internally at FutureHouse, we have a slightly different set of tools. We're trying to get some of them, like citation traversal, into this repo. However, we have APIs and licenses to access research papers that we cannot share openly. Similarly, in our research papers' results we do not start with the known relevant PDFs. Our agent has to identify them using keyword search over all papers, rather than just a subset. We're gradually aligning these two versions of PaperQA, but until there is an open-source way to freely access papers (even just open source papers) you will need to provide PDFs yourself.

### How is this different from LlamaIndex or LangChain?

[LangChain](https://github.com/langchain-ai/langchain) and [LlamaIndex](https://github.com/run-llama/llama_index) are both frameworks for working with LLM applications, with abstractions made for agentic workflows and retrieval augmented generation.

Over time, the PaperQA team over time chose to become framework-agnostic, instead outsourcing LLM drivers to [LiteLLM](https://docs.litellm.ai/docs/) and no framework besides Pydantic for its tools. PaperQA focuses on scientific papers and their metadata.

PaperQA can be reimplemented using either LlamaIndex or LangChain. For example, our `GatherEvidence` tool can be reimplemented as a retriever with an LLM-based re-ranking and contextual summary. There is similar work with the tree response method in LlamaIndex.

### Can I save or load?

The `Docs` class can be pickled and unpickled. This is useful if you want to save the embeddings of the documents and then load them later.

```python
import pickle

# save
with open("my_docs.pkl", "wb") as f:
    pickle.dump(docs, f)

# load
with open("my_docs.pkl", "rb") as f:
    docs = pickle.load(f)
```

## Reproduction

Contained in [docs/2024-10-16\_litqa2-splits.json5](https://github.com/Future-House/paper-qa/blob/main/docs/2024-10-16_litqa2-splits.json5) are the question IDs used in train, evaluation, and test splits, as well as paper DOIs used to build the splits' indexes.

* Train and eval splits: question IDs come from [LAB-Bench's LitQA2 question IDs](https://github.com/Future-House/LAB-Bench/blob/main/LitQA2/litqa-v2-public.jsonl).
* Test split: questions IDs come from [aviary-paper-data's LitQA2 question IDs](https://huggingface.co/datasets/futurehouse/aviary-paper-data).

There are multiple papers slowly building PaperQA, shown below in [Citation](#citation). To reproduce:

* `skarlinski2024language`: train and eval splits are applicable. The test split remains held out.
* `narayanan2024aviarytraininglanguageagents`: train, eval, and test splits are applicable.

Example on how to use LitQA for evaluation can be found in [aviary.litqa](https://github.com/Future-House/aviary/tree/main/packages/litqa#running-litqa).

## Citation

Please read and cite the following papers if you use this software:

```bibtex
@article{narayanan2024aviarytraininglanguageagents,
      title = {Aviary: training language agents on challenging scientific tasks},
      author = {
      Siddharth Narayanan and
 James D. Braza and
 Ryan-Rhys Griffiths and
 Manu Ponnapati and
 Albert Bou and
 Jon Laurent and
 Ori Kabeli and
 Geemi Wellawatte and
 Sam Cox and
 Samuel G. Rodriques and
 Andrew D. White},
      journal = {arXiv preprent arXiv:2412.21154},
      year = {2024},
      url = {https://doi.org/10.48550/arXiv.2412.21154},
}
```

```bibtex
@article{skarlinski2024language,
    title = {Language agents achieve superhuman synthesis of scientific knowledge},
    author = {
    Michael D. Skarlinski and
 Sam Cox and
 Jon M. Laurent and
 James D. Braza and
 Michaela Hinks and
 Michael J. Hammerling and
 Manvitha Ponnapati and
 Samuel G. Rodriques and
 Andrew D. White},
    journal = {arXiv preprent arXiv:2409.13740},
    year = {2024},
    url = {https://doi.org/10.48550/arXiv.2409.13740}
}
```

```bibtex
@article{lala2023paperqa,
    title = {PaperQA: Retrieval-Augmented Generative Agent for Scientific Research},
    author = {
    Jakub Lála and
 Odhran O'Donoghue and
 Aleksandar Shtedritski and
 Sam Cox and
 Samuel G. Rodriques and
 Andrew D. White},
    journal = {arXiv preprint arXiv:2312.07559},
    year = {2023},
    url = {https://doi.org/10.48550/arXiv.2312.07559}
}
```


# Contributing to PaperQA

Thank you for your interest in contributing to PaperQA! Here are some guidelines to help you get started.

## Setting up the development environment

We use [`uv`](https://github.com/astral-sh/uv) for our local development.

1. Install `uv` by following the instructions on the [uv website](https://astral.sh/uv/).
2. Run the following command to install all dependencies and set up the development environment:

   ```bash
   uv sync
   ```

## Installing the package for development

If you prefer to use `pip` for installing the package in development mode, you can do so by running:

```bash
pip install -e ".[dev]"
```

Where the `dev` extra includes development dependencies such as `pytest`.

## Running tests and other tooling

Use the following commands:

* Run tests (requires an OpenAI key in your environment)

  ```bash
  pytest
  # or for multiprocessing based parallelism
  pytest -n auto
  ```
* Run `pre-commit` for formatting and type checking

  ```bash
  pre-commit run --all-files
  ```
* Run `mypy`, `refurb`, or `pylint` directly:

  ```bash
  mypy paperqa
  # or
  refurb paperqa
  # or
  pylint paperqa
  ```

See our GitHub Actions [`tests.yml`](https://github.com/Future-House/paper-qa/blob/main/.github/workflows/tests.yml) for further reference.

## Using `pytest-recording` and VCR cassettes

We use the [`pytest-recording`](https://github.com/kiwicom/pytest-recording) plugin to create VCR cassettes to cache HTTP requests, making our unit tests more deterministic.

To record a new VCR cassette:

```bash
uv run pytest --record-mode=once tests/desired_test_module.py
```

And the new cassette(s) should appear in [`tests/cassettes`](https://github.com/Future-House/paper-qa/blob/main/tests/cassettes/README.md).

Our configuration for `pytest-recording` can be found in [`tests/conftest.py`](https://github.com/Future-House/paper-qa/blob/main/tests/conftest.py). This includes header removals (e.g. OpenAI `authorization` key) from responses to ensure sensitive information is excluded from the cassettes.

Please ensure cassettes are less than 1 MB to keep tests loading quickly.

Happy coding!


# tests


# SF Districts in the style of Andy Warhol

<img src="https://upload.wikimedia.org/wikipedia/commons/thumb/1/12/Wikimedia_logo_text_RGB.svg/330px-Wikimedia_logo_text_RGB.svg.png" alt="Wikimedia logo" height="100">

![Map of SF districts](/files/DuXrqQKcSno7YtxieBgC)

Text under image 1.

| Col1  | Col2  |
| ----- | ----- |
| Val11 | Val12 |
| Val21 | Val11 |

Text under table 1.

Inline LaTeX: $E = mc^2$

Block LaTeX:

$$
x + n + a = \sqrt{ax + (n + a)^2 + x \sqrt{a (x + n) + (n + a)^2 + (x + n) \sqrt{\dots}}}
$$

<img src="https://upload.wikimedia.org/wikipedia/commons/thumb/1/12/Wikimedia_logo_text_RGB.svg/330px-Wikimedia_logo_text_RGB.svg.png" alt="Wikimedia logo" height="100">

![Map of SF districts](/files/DuXrqQKcSno7YtxieBgC)

Text under image 2.

| Col1  | Col2  |
| ----- | ----- |
| Val11 | Val12 |
| Val21 | Val11 |

Text under table 2.

Inline LaTeX: $E = mc^2$

Block LaTeX:

$$
x + n + a = \sqrt{ax + (n + a)^2 + x \sqrt{a (x + n) + (n + a)^2 + (x + n) \sqrt{\dots}}}
$$

<img src="https://upload.wikimedia.org/wikipedia/commons/thumb/1/12/Wikimedia_logo_text_RGB.svg/330px-Wikimedia_logo_text_RGB.svg.png" alt="Wikimedia logo" height="100">

![Map of SF districts](/files/DuXrqQKcSno7YtxieBgC)

Text under image 3.

| Col1  | Col2  |
| ----- | ----- |
| Val11 | Val12 |
| Val21 | Val11 |

Text under table 3.

Inline LaTeX: $E = mc^2$

Block LaTeX:

$$
x + n + a = \sqrt{ax + (n + a)^2 + x \sqrt{a (x + n) + (n + a)^2 + (x + n) \sqrt{\dots}}}
$$

<img src="https://upload.wikimedia.org/wikipedia/commons/thumb/1/12/Wikimedia_logo_text_RGB.svg/330px-Wikimedia_logo_text_RGB.svg.png" alt="Wikimedia logo" height="100">

![Map of SF districts](/files/DuXrqQKcSno7YtxieBgC)

Text under image 4.

| Col1  | Col2  |
| ----- | ----- |
| Val11 | Val12 |
| Val21 | Val11 |

Text under table 4.

Inline LaTeX: $E = mc^2$

Block LaTeX:

$$
x + n + a = \sqrt{ax + (n + a)^2 + x \sqrt{a (x + n) + (n + a)^2 + (x + n) \sqrt{\dots}}}
$$

<img src="https://upload.wikimedia.org/wikipedia/commons/thumb/1/12/Wikimedia_logo_text_RGB.svg/330px-Wikimedia_logo_text_RGB.svg.png" alt="Wikimedia logo" height="100">

![Map of SF districts](/files/DuXrqQKcSno7YtxieBgC)

Text under image 5.

| Col1  | Col2  |
| ----- | ----- |
| Val11 | Val12 |
| Val21 | Val11 |

Text under table 5.

Inline LaTeX: $E = mc^2$

Block LaTeX:

$$
x + n + a = \sqrt{ax + (n + a)^2 + x \sqrt{a (x + n) + (n + a)^2 + (x + n) \sqrt{\dots}}}
$$


# stub\_data


# Gravity hill

> "Magnetic hill" and "Mystery hill" redirect here. For other uses, see [Magnetic Hill (disambiguation)](https://en.wikipedia.org/wiki/Magnetic_Hill_\(disambiguation\)) and [Mystery Hill (disambiguation)](https://en.wikipedia.org/wiki/Mystery_Hill).

A **gravity hill**, also known as a **magnetic hill**, **mystery hill**, **mystery spot**, **gravity road**, or **anti-gravity hill**, is a place where the layout of the surrounding land produces an [illusion](https://en.wikipedia.org/wiki/Illusion), making a slight downhill slope appear to be an uphill slope. Thus, a car left out of gear will appear to be rolling uphill against [gravity](https://en.wikipedia.org/wiki/Gravity).

Although the slope of gravity hills is an illusion, sites are often accompanied by claims that magnetic or supernatural forces are at work. The most important factor contributing to the illusion is a completely or mostly obstructed horizon. Without a horizon, it becomes difficult for a person to judge the slope of a surface, as a reliable reference point is missing, and misleading visual cues can adversely affect the sense of balance. Objects which one would normally assume to be more or less perpendicular to the ground, such as trees, may be leaning, offsetting the visual reference.

A 2003 study looked into how the absence of a horizon can skew the perspective on gravity hills, by recreating a number of antigravity places in the lab to see how volunteers would react. In conclusion, researchers from the Universities of Padua and Pavia in Italy found that without a true horizon in sight, the human brain could be tricked by common landmarks such as trees and signs.

The illusion is similar to the [Ames room](https://en.wikipedia.org/wiki/Ames_room), in which objects can also appear to roll against gravity.

The opposite phenomenon—an uphill road that appears flat—is known in [bicycle racing](https://en.wikipedia.org/wiki/Cycle_sport) as a ["false flat"](https://en.wikipedia.org/wiki/Glossary_of_cycling#F).

## See also

* [List of gravity hills](https://en.wikipedia.org/wiki/List_of_gravity_hills)
* [The Crooked House](https://en.wikipedia.org/wiki/The_Crooked_House) – a pub (now demolished) with an internal gravity hill illusion.

## References

## External links


# docs


# tutorials


# PaperQA2 for Clinical Trials

PaperQA2 now natively supports querying clinical trials in addition to any documents supplied by the user. It uses a new tool, the aptly named `clinical_trials_search` tool. Users don't have to provide any clinical trials to the tool itself, it uses the `clinicaltrials.gov` API to retrieve them on the fly. As of January 2025, the tool is not enabled by default, but it's easy to configure. Here's an example where we query only clinical trials, without using any documents:

```python
from paperqa import Settings, agent_query

answer_response = await agent_query(
    query="What drugs have been found to effectively treat Ulcerative Colitis?",
    settings=Settings.from_name("search_only_clinical_trials"),
)

print(answer_response.session.answer)
```

### Output

```
Several drugs have been found to effectively treat Ulcerative Colitis (UC),
targeting different mechanisms of the disease.

Golimumab, a tumor necrosis factor (TNF) inhibitor marketed as Simponi®, has demonstrated efficacy
in treating moderate-to-severe UC. Administered subcutaneously, it was shown to maintain clinical
response through Week 54 in patients, as assessed by the Partial Mayo Score (NCT02092285).

Mesalazine, an anti-inflammatory drug, is commonly used for UC treatment. In a study comparing
mesalazine enemas to faecal microbiota transplantation (FMT) for left-sided UC,
mesalazine enemas (4g daily) were effective in inducing clinical remission (Mayo score ≤ 2) (NCT03104036).

Antibiotics have also shown potential in UC management. A combination of doxycycline,
amoxicillin, and metronidazole induced remission in 60-70% of patients with moderate-to-severe
UC in prior studies. These antibiotics are thought to alter gut microbiota, reducing pathobionts
 and promoting beneficial bacteria (NCT02217722, NCT03986996).

Roflumilast, a phosphodiesterase-4 (PDE4) inhibitor, is being investigated for mild-to-moderate UC.
Preliminary findings suggest it may improve disease severity and biochemical markers when
added to conventional treatments (NCT05684484).

These treatments highlight diverse therapeutic approaches, including immunosuppression,
microbiota modulation, and anti-inflammatory mechanisms.
```

You can see the in-line citations for each clinical trial used as a response for each query. If you'd like to see more data on the specific contexts that were used to answer the query:

```python
print(answer_response.session.contexts)
```

```
[Context(context='The excerpt mentions that a search on ClinicalTrials.gov for clinical trials related to drugs
treating Ulcerative Colitis yielded 689 trials. However, it does not provide specific information about which
drugs have been found effective for treating Ulcerative Colitis.', text=Text(text='', name=...
```

Using `Settings.from_name('search_only_clinical_trials')` is a shortcut, but note that you can easily add `clinical_trial_search` into any custom `Settings` by just explicitly naming it as a tool:

```python
from pathlib import Path
from paperqa import Settings, agent_query, AgentSetting
from paperqa.agents.tools import DEFAULT_TOOL_NAMES

# you can start with the default list of PaperQA tools
print(DEFAULT_TOOL_NAMES)
# >>> ['paper_search', 'gather_evidence', 'gen_answer', 'reset', 'complete'],

# we can start with a directory with a potentially useful paper in it
print(list(Path("my_papers").iterdir()))

# now let's query using standard tools + clinical_trials
answer_response = await agent_query(
    query="What drugs have been found to effectively treat Ulcerative Colitis?",
    settings=Settings(
        paper_directory="my_papers",
        agent={"tool_names": DEFAULT_TOOL_NAMES + ["clinical_trials_search"]},
    ),
)

# let's check out the formatted answer (with references included)
print(answer_response.session.formatted_answer)
```

```
Question: What drugs have been found to effectively treat Ulcerative Colitis?

Several drugs have been found effective in treating Ulcerative Colitis (UC), with treatment
strategies varying based on disease severity and extent. For mild-to-moderate UC, 5-aminosalicylic
 acid (5-ASA) is the first-line therapy. Topical 5-ASA, such as mesalazine suppositories (1 g/day),
 is effective for proctitis or distal colitis, inducing remission in 31-80% of patients. Oral mesalazine
 at higher doses (e.g., 4.8 g/day) can accelerate clinical improvement in more extensive disease
 (meier2011currenttreatmentof pages 1-2; meier2011currenttreatmentof pages 3-4).

For moderate-to-severe cases, corticosteroids are commonly used. Oral steroids like prednisolone
(40-60 mg/day) or intravenous steroids such as methylprednisolone (60 mg/day) and hydrocortisone
(400 mg/day) are standard for inducing remission (meier2011currenttreatmentof pages 3-4). Tumor
necrosis factor (TNF)-α blockers, such as infliximab, are effective for steroid-refractory cases
(meier2011currenttreatmentof pages 2-3; meier2011currenttreatmentof pages 3-4).

Immunosuppressive agents, including azathioprine and 6-mercaptopurine, are used for maintenance
therapy in steroid-dependent or refractory cases (meier2011currenttreatmentof pages 2-3;
meier2011currenttreatmentof pages 3-4). Antibiotics, such as combinations of penicillin,
tetracycline, and metronidazole, have shown promise in altering the microbiota and inducing
remission in some patients, though their efficacy varies (NCT02217722).

References

1. (meier2011currenttreatmentof pages 2-3): Johannes Meier and Andreas Sturm. Current treatment
of ulcerative colitis. World journal of gastroenterology, 17 27:3204-12, 2011.
URL: https://doi.org/10.3748/wjg.v17.i27.3204, doi:10.3748/wjg.v17.i27.3204.

2. (meier2011currenttreatmentof pages 3-4): Johannes Meier and Andreas Sturm. Current treatment
of ulcerative colitis. World journal of gastroenterology, 17 27:3204-12, 2011. URL:
https://doi.org/10.3748/wjg.v17.i27.3204, doi:10.3748/wjg.v17.i27.3204.

3. (NCT02217722): Prof. Arie Levine. Use of the Ulcerative Colitis Diet for Induction of
Remission. Prof. Arie Levine. 2014. ClinicalTrials.gov Identifier: NCT02217722

4. (meier2011currenttreatmentof pages 1-2): Johannes Meier and Andreas Sturm. Current
treatment of ulcerative colitis. World journal of gastroenterology, 17 27:3204-12, 2011.
 URL: https://doi.org/10.3748/wjg.v17.i27.3204, doi:10.3748/wjg.v17.i27.3204.
```

We now see both papers and clinical trials cited in our response. For convenience, we have a `Settings.from_name` that works as well:

```python
from paperqa import Settings, agent_query

answer_response = await agent_query(
    query="What drugs have been found to effectively treat Ulcerative Colitis?",
    settings=Settings.from_name("clinical_trials"),
)
```

And, this works with the `pqa` cli as well:

```bash
>>> pqa --settings 'search_only_clinical_trials' ask 'what is Ibuprofen effective at treating?'
```

```
...
[13:29:50] Completing 'what is Ibuprofen effective at treating?' as 'certain'.
        Answer: Ibuprofen is a non-steroidal anti-inflammatory drug (NSAID) effective
        in treating various conditions, including pain, inflammation, and fever.
        It is widely used for tension-type
        headaches, with studies showing that ibuprofen sodium provides significant
        pain relief and reduces pain intensity compared to standard ibuprofen and placebo
        over a 3-hour period (NCT01362491).
        Intravenous ibuprofen is effective in managing postoperative pain, particularly
        in orthopedic surgeries, and helps control the inflammatory process. When combined
        with opioids, it reduces opioid
        consumption and associated side effects, making it a key component of
        multimodal analgesia (NCT05401916, NCT01773005).

        Ibuprofen is also effective in pediatric populations as a first-line
        anti-inflammatory and antipyretic agent due to its relatively
        low adverse effects compared to other NSAIDs (NCT01478022).
        Additionally, it has been studied for its potential use in managing
        chronic periodontitis through subgingival irrigation with a 2% ibuprofen
        mouthwash, which reduces periodontal pocket depth and
        bleeding on probing, improving periodontal health (NCT02538237).

        These findings highlight ibuprofen's versatility in treating pain, inflammation,
        fever, and specific conditions like tension headaches, postoperative pain, and periodontal diseases.
```


# Measuring PaperQA2 with LFRQA

> This tutorial is available as a Jupyter notebook [here](/paperqa/docs/tutorials/running_on_lfrqa)

## Overview

The **LFRQA dataset** was introduced in the paper [*RAG-QA Arena: Evaluating Domain Robustness for Long-Form Retrieval-Augmented Question Answering*](https://arxiv.org/pdf/2407.13998). It features **1,404 science questions** (along with other categories) that have been human-annotated with answers. This tutorial walks through the process of setting up the dataset for use and benchmarking.

## Download the Annotations

First, we need to obtain the annotated dataset from the official repository:

```python
# Create a new directory for the dataset
!mkdir -p data/rag-qa-benchmarking

# Get the annotated questions
!curl https://raw.githubusercontent.com/awslabs/rag-qa-arena/refs/heads/main/data/\
annotations_science_with_citation.jsonl \
-o data/rag-qa-benchmarking/annotations_science_with_citation.jsonl
```

## Download the Robust-QA Documents

LFRQA is built upon **Robust-QA**, so we must download the relevant documents:

```python
# Download the Lotte dataset, which includes the required documents
!curl https://downloads.cs.stanford.edu/nlp/data/colbert/colbertv2/lotte.tar.gz --output lotte.tar.gz

# Extract the dataset
!tar -xvzf lotte.tar.gz

# Move the science test collection to our dataset folder
!cp lotte/science/test/collection.tsv ./data/rag-qa-benchmarking/science_test_collection.tsv

# Clean up unnecessary files
!rm lotte.tar.gz
!rm -rf lotte
```

For more details, refer to the original paper: [*RAG-QA Arena: Evaluating Domain Robustness for Long-Form Retrieval-Augmented Question Answering*](https://arxiv.org/pdf/2407.13998).

## Load the Data

We now load the documents into a pandas dataframe:

```python
import os

import pandas as pd

# Load questions and answers dataset
rag_qa_benchmarking_dir = os.path.join("data", "rag-qa-benchmarking")

# Load documents dataset
lfrqa_docs_df = pd.read_csv(
    os.path.join(rag_qa_benchmarking_dir, "science_test_collection.tsv"),
    sep="\t",
    names=["doc_id", "doc_text"],
)
```

## Select the Documents to Use

RobustQA consists on 1.7M documents. Hence, it takes around 3 hours to build the whole index.

To run a test, we can use 1% of the dataset. This will be accomplished by selecting the first 1% available documents and the questions referent to these documents.

```python
proportion_to_use = 1 / 100
amount_of_docs_to_use = int(len(lfrqa_docs_df) * proportion_to_use)
print(f"Using {amount_of_docs_to_use} out of {len(lfrqa_docs_df)} documents")
```

## Prepare the Document Files

We now create the document directory and store each document as a separate text file, so that paperqa can build the index.

```python
partial_docs = lfrqa_docs_df.head(amount_of_docs_to_use)
lfrqa_directory = os.path.join(rag_qa_benchmarking_dir, "lfrqa")
os.makedirs(
    os.path.join(lfrqa_directory, "science_docs_for_paperqa", "files"), exist_ok=True
)

for i, row in partial_docs.iterrows():
    doc_id = row["doc_id"]
    doc_text = row["doc_text"]

    with open(
        os.path.join(
            lfrqa_directory, "science_docs_for_paperqa", "files", f"{doc_id}.txt"
        ),
        "w",
        encoding="utf-8",
    ) as f:
        f.write(doc_text)

    if i % int(len(partial_docs) * 0.05) == 0:
        progress = (i + 1) / len(partial_docs)
        print(f"Progress: {progress:.2%}")
```

## Create the Manifest File

The **manifest file** keeps track of document metadata for the dataset. We need to fill some fields so that paperqa doesn’t try to get metadata using llm calls. This will make the indexing process faster.

```python
manifest = partial_docs.copy()
manifest["file_location"] = manifest["doc_id"].apply(lambda x: f"files/{x}.txt")
manifest["doi"] = ""
manifest["title"] = manifest["doc_id"]
manifest["key"] = manifest["doc_id"]
manifest["docname"] = manifest["doc_id"]
manifest["citation"] = "_"
manifest = manifest.drop(columns=["doc_id", "doc_text"])
manifest.to_csv(
    os.path.join(lfrqa_directory, "science_docs_for_paperqa", "manifest.csv"),
    index=False,
)
```

## Filter and Save Questions

Finally, we load the questions and filter them to ensure we only include questions that reference the selected documents:

```python
questions_df = pd.read_json(
    os.path.join(rag_qa_benchmarking_dir, "annotations_science_with_citation.jsonl"),
    lines=True,
)
partial_questions = questions_df[
    questions_df.gold_doc_ids.apply(
        lambda ids: all(_id < amount_of_docs_to_use for _id in ids)
    )
]
partial_questions.to_csv(
    os.path.join(lfrqa_directory, "questions.csv"),
    index=False,
)

print("Using", len(partial_questions), "questions")
```

## Install paperqa

From now on, we will be using the paperqa library, so we need to install it:

```python
!pip install paper-qa
```

## Index the Documents

Now we will build an index for the LFRQA documents. The index is a **Tantivy index**, which is a fast, full-text search engine library written in Rust. Tantivy is designed to handle large datasets efficiently, making it ideal for searching through a vast collection of papers or documents.

Feel free to adjust the concurrency settings as you like. Because we defined a manifest, we don’t need any API keys for building this index because we don't discern any citation metadata, but you do need LLM API keys to answer questions.

Remember that this process is quick for small portions of the dataset, but can take around 3 hours for the whole dataset.

```python
import nest_asyncio

nest_asyncio.apply()
```

We add the line above to handle async code within a notebook.

However, to improve compatibility and speed up the indexing process, we strongly recommend running the following code in a separate `.py` file

```python
import os

from paperqa import Settings
from paperqa.agents import build_index
from paperqa.settings import AgentSettings, IndexSettings, ParsingSettings

settings = Settings(
    agent=AgentSettings(
        index=IndexSettings(
            name="lfrqa_science_index",
            paper_directory=os.path.join(
                "data", "rag-qa-benchmarking", "lfrqa", "science_docs_for_paperqa"
            ),
            index_directory=os.path.join(
                "data", "rag-qa-benchmarking", "lfrqa", "science_docs_for_paperqa_index"
            ),
            manifest_file="manifest.csv",
            concurrency=10_000,
            batch_size=10_000,
        )
    ),
    parsing=ParsingSettings(
        use_doc_details=False,
        defer_embedding=True,
    ),
)

build_index(settings=settings)
```

After this runs, you will have an index ready to use!

## Benchmark!

After you have built the index, you are ready to run the benchmark. We advice running this in a separate `.py` file.

To run this, you will need to have the [`ldp`](https://github.com/Future-House/ldp) and [`fhaviary[lfrqa]`](https://github.com/Future-House/aviary/blob/main/packages/lfrqa/README.md#installation) packages installed.

```python
!pip install ldp "fhaviary[lfrqa]"
```

```python
import asyncio
import json
import logging
import os

import pandas as pd
from aviary.envs.lfrqa import LFRQAQuestion, LFRQATaskDataset
from ldp.agent import SimpleAgent
from ldp.alg.runners import Evaluator, EvaluatorConfig

from paperqa import Settings
from paperqa.settings import AgentSettings, IndexSettings

logging.basicConfig(level=logging.ERROR)

log_results_dir = os.path.join("data", "rag-qa-benchmarking", "results")
os.makedirs(log_results_dir, exist_ok=True)


async def log_evaluation_to_json(  # noqa: RUF029
    lfrqa_question_evaluation: dict,
) -> None:
    json_path = os.path.join(
        log_results_dir, f"{lfrqa_question_evaluation['qid']}.json"
    )
    with open(json_path, "w") as f:  # noqa: ASYNC230
        json.dump(lfrqa_question_evaluation, f, indent=2)


async def evaluate() -> None:
    settings = Settings(
        agent=AgentSettings(
            index=IndexSettings(
                name="lfrqa_science_index",
                paper_directory=os.path.join(
                    "data", "rag-qa-benchmarking", "lfrqa", "science_docs_for_paperqa"
                ),
                index_directory=os.path.join(
                    "data",
                    "rag-qa-benchmarking",
                    "lfrqa",
                    "science_docs_for_paperqa_index",
                ),
            )
        )
    )

    data: list[LFRQAQuestion] = [
        LFRQAQuestion(**row)
        for row in pd.read_csv(
            os.path.join("data", "rag-qa-benchmarking", "lfrqa", "questions.csv")
        )[["qid", "question", "answer", "gold_doc_ids"]].to_dict(orient="records")
    ]

    dataset = LFRQATaskDataset(
        data=data,
        settings=settings,
        evaluation_callback=log_evaluation_to_json,
    )

    evaluator = Evaluator(
        config=EvaluatorConfig(batch_size=3),
        agent=SimpleAgent(),
        dataset=dataset,
    )
    await evaluator.evaluate()


if __name__ == "__main__":
    asyncio.run(evaluate())
```

After running this, you can find the results in the `data/rag-qa-benchmarking/results` folder. Here is an example of how to read them:

```python
import glob

json_files = glob.glob(os.path.join(rag_qa_benchmarking_dir, "results", "*.json"))

data = []
for file in json_files:
    with open(file) as f:
        json_data = json.load(f)
        json_data["qid"] = file.split("/")[-1].replace(".json", "")
        data.append(json_data)

results_df = pd.DataFrame(data).set_index("qid")
results_df["winner"].value_counts(normalize=True)
```


# settings\_tutorial

## Setup

> This tutorial is available as a Jupyter notebook [here](https://github.com/Future-House/paper-qa/blob/main/docs/tutorials/settings_tutorial.ipynb).

This tutorial aims to show how to use the `Settings` class to configure `PaperQA`. Firstly, we will be using `OpenAI` and `Anthropic` models, so we need to set the `OPENAI_API_KEY` and `ANTHROPIC_API_KEY` environment variables. We will use both models to make it clear when `paperqa` agent is using either one or the other. We use `python-dotenv` to load the environment variables from a `.env` file. Hence, our first step is to create a `.env` file and install the required packages.

```python
# fmt: off
# Create .env file with OpenAI API and Anthropic API keys
# Replace <your-openai-api-key> and <your-anthropic-api-key> with your actual API keys
!echo "OPENAI_API_KEY=<your-openai-api-key>" > .env # fmt: skip
!echo "ANTHROPIC_API_KEY=<your-anthropic-api-key>" >> .env # fmt: skip

!uv pip install -q nest-asyncio python-dotenv aiohttp fhlmi "paper-qa[local]"
# fmt: on
```

```python
import os

import aiohttp
import nest_asyncio
from dotenv import load_dotenv

nest_asyncio.apply()
load_dotenv(".env")
```

```python
print("You have set the following environment variables:")
print(
    f"OPENAI_API_KEY:    {'is set' if os.environ['OPENAI_API_KEY'] else 'is not set'}"
)
print(
    f"ANTHROPIC_API_KEY: {'is set' if os.environ['ANTHROPIC_API_KEY'] else 'is not set'}"
)
```

We will use the `lmi` package to get the model names and the `.papers` directory to save documents we will use.

```python
from lmi import CommonLLMNames

llm_openai = CommonLLMNames.OPENAI_TEST.value
llm_anthropic = CommonLLMNames.ANTHROPIC_TEST.value

# Create the `papers` directory if it doesn't exist
os.makedirs("papers", exist_ok=True)

# Download the paper from arXiv and save it to the `papers` directory
url = "https://arxiv.org/pdf/2407.01603"
async with aiohttp.ClientSession() as session, session.get(url, timeout=60) as response:
    content = await response.read()
    with open("papers/2407.01603.pdf", "wb") as f:
        f.write(content)
```

The `Settings` class is used to configure the PaperQA settings. Official documentation can be found [here](https://github.com/Future-House/paper-qa?tab=readme-ov-file#settings-cheatsheet) and the open source code can be found [here](https://github.com/Future-House/paper-qa/blob/main/src/paperqa/settings.py).

Here is a basic example of how to use the `Settings` class. We will be unnecessarily verbose for the sake of clarity. Please notice that most of the settings are optional and the defaults are good for most cases. Refer to the [descriptions of each setting](https://github.com/Future-House/paper-qa/blob/main/src/paperqa/settings.py) for more information.

Within this `Settings` object, I'd like to discuss specifically how the llms are configured and how `paperqa` looks for papers.

A common source of confusion is that multiple `llms` are used in paperqa. We have `llm`, `summary_llm`, `agent_llm`, and `embedding`. Hence, if `llm` is set to an `Anthropic` model, `summary_llm` and `agent_llm` will still require a `OPENAI_API_KEY`, since `OpenAI` models are the default.

Among the objects that use `llms` in `paperqa`, we have `llm`, `summary_llm`, `agent_llm`, and `embedding`:

* `llm`: Main LLM used by the agent to reason about the question, extract metadata from documents, etc.
* `summary_llm`: LLM used to summarize the papers.
* `agent_llm`: LLM used to answer questions and select tools.
* `embedding`: Embedding model used to embed the papers.

Let's see some examples around this concept. First, we define the settings with `llm` set to an `OpenAI` model. Please notice this is not an complete list of settings. But take your time to read through this `Settings` class and all customization that can be done.

```python
import pathlib

from paperqa.prompts import (
    CONTEXT_INNER_PROMPT,
    CONTEXT_OUTER_PROMPT,
    citation_prompt,
    default_system_prompt,
    env_reset_prompt,
    env_system_prompt,
    qa_prompt,
    select_paper_prompt,
    structured_citation_prompt,
    summary_json_prompt,
    summary_json_system_prompt,
    summary_prompt,
)
from paperqa.settings import (
    AgentSettings,
    AnswerSettings,
    IndexSettings,
    ParsingSettings,
    PromptSettings,
    Settings,
)

settings = Settings(
    llm=llm_openai,
    llm_config={
        "model_list": [
            {
                "model_name": llm_openai,
                "litellm_params": {
                    "model": llm_openai,
                    "temperature": 0.1,
                    "max_tokens": 4096,
                },
            }
        ],
        "rate_limit": {
            llm_openai: "30000 per 1 minute",
        },
    },
    summary_llm=llm_openai,
    summary_llm_config={
        "rate_limit": {
            llm_openai: "30000 per 1 minute",
        },
    },
    embedding="text-embedding-3-small",
    embedding_config={},
    temperature=0.1,
    batch_size=1,
    verbosity=1,
    answer=AnswerSettings(
        evidence_k=10,
        evidence_retrieval=True,
        evidence_summary_length="about 100 words",
        evidence_skip_summary=False,
        answer_max_sources=5,
        max_answer_attempts=None,
        answer_length="about 200 words, but can be longer",
        max_concurrent_requests=10,
    ),
    parsing=ParsingSettings(
        reader_config={"chunk_chars": 5000, "overlap": 250},
        citation_prompt=citation_prompt,
        structured_citation_prompt=structured_citation_prompt,
    ),
    prompts=PromptSettings(
        summary=summary_prompt,
        qa=qa_prompt,
        select=select_paper_prompt,
        pre=None,
        post=None,
        system=default_system_prompt,
        use_json=True,
        summary_json=summary_json_prompt,
        summary_json_system=summary_json_system_prompt,
        context_outer=CONTEXT_OUTER_PROMPT,
        context_inner=CONTEXT_INNER_PROMPT,
    ),
    agent=AgentSettings(
        agent_llm=llm_openai,
        agent_llm_config={
            "model_list": [
                {
                    "model_name": llm_openai,
                    "litellm_params": {
                        "model": llm_openai,
                    },
                }
            ],
            "rate_limit": {
                llm_openai: "30000 per 1 minute",
            },
        },
        agent_prompt=env_reset_prompt,
        agent_system_prompt=env_system_prompt,
        search_count=8,
        index=IndexSettings(
            paper_directory=pathlib.Path.cwd().joinpath("papers"),
            manifest_file=None,
            index_directory=pathlib.Path.cwd().joinpath("papers/index"),
        ),
    ),
)
```

As it is evident, `Paperqa` is absolutely customizable. And here we reinterate that despite this possible fine customization, the defaults are good for most cases. Although, the user is welcome to explore the settings and customize the `paperqa` to their needs.

We also set settings.verbosity to 1, which will print the agent configuration. Feel free to set it to 0 to silence the logging after your first run.

```python
from paperqa import ask

response = ask(
    "What are the most relevant language models used for chemistry?", settings=settings
)
```

Which probably worked fine. Let's now try to remove `OPENAI_API_KEY` and run again the same question with the same settings.

```python
os.environ["OPENAI_API_KEY"] = ""
print("You have set the following environment variables:")
print(
    f"OPENAI_API_KEY:    {'is set' if os.environ['OPENAI_API_KEY'] else 'is not set'}"
)
print(
    f"ANTHROPIC_API_KEY: {'is set' if os.environ['ANTHROPIC_API_KEY'] else 'is not set'}"
)
```

```python
response = ask(
    "What are the most relevant language models used for chemistry?", settings=settings
)
```

It would obviously fail. We don't have a valid `OPENAI_API_KEY`, so the agent will not be able to use `OpenAI` models. Let's change it to an `Anthropic` model and see if it works.

```python
settings.llm = llm_anthropic
settings.llm_config = {
    "model_list": [
        {
            "model_name": llm_anthropic,
            "litellm_params": {
                "model": llm_anthropic,
                "temperature": 0.1,
                "max_tokens": 512,
            },
        }
    ],
    "rate_limit": {
        llm_anthropic: "30000 per 1 minute",
    },
}
settings.summary_llm = llm_anthropic
settings.summary_llm_config = {
    "rate_limit": {
        llm_anthropic: "30000 per 1 minute",
    },
}
settings.agent = AgentSettings(
    agent_llm=llm_anthropic,
    agent_llm_config={
        "rate_limit": {
            llm_anthropic: "30000 per 1 minute",
        },
    },
    index=IndexSettings(
        paper_directory=pathlib.Path.cwd().joinpath("papers"),
        manifest_file=None,
        index_directory=pathlib.Path.cwd().joinpath("papers/index"),
    ),
)
settings.embedding = "st-multi-qa-MiniLM-L6-cos-v1"
response = ask(
    "What are the most relevant language models used for chemistry?", settings=settings
)
```

Now the agent is able to use `Anthropic` models only and although we don't have a valid `OPENAI_API_KEY`, the question is answered because the agent will not use `OpenAI` models. See that we also changed the `embedding` because it was using `text-embedding-3-small` by default, which is a `OpenAI` model. `Paperqa` implements a few embedding models. Please refer to the [documentation](https://github.com/Future-House/paper-qa?tab=readme-ov-file#embedding-model) for more information.

In addition, notice that this is a very verbose example for the sake of clarity. We could have just set only the llms names and used default settings for the rest:

```python
llm_anthropic_config = {
    "model_list": [{
            "model_name": llm_anthropic,
    }]
}

settings.llm = llm_anthropic
settings.llm_config = llm_anthropic_config
settings.summary_llm = llm_anthropic
settings.summary_llm_config = llm_anthropic_config
settings.agent = AgentSettings(
    agent_llm=llm_anthropic,
    agent_llm_config=llm_anthropic_config,
    index=IndexSettings(
        paper_directory=pathlib.Path.cwd().joinpath("papers"),
        manifest_file=None,
        index_directory=pathlib.Path.cwd().joinpath("papers/index"),
    ),
)
settings.embedding = "st-multi-qa-MiniLM-L6-cos-v1"
```

## The output

`Paperqa` returns a `PQASession` object, which contains not only the answer but also all the information gatheres to answer the questions. We recommend printing the `PQASession` object (`print(response.session)`) to understand the information it contains. Let's check the `PQASession` object:

```python
print(response.session)
```

```python
print("Let's examine the PQASession object returned by paperqa:\n")

print(f"Status: {response.status.value}")

print("1. Question asked:")
print(f"{response.session.question}\n")

print("2. Answer provided:")
print(f"{response.session.answer}\n")
```

In addition to the answer, the `PQASession` object contains all the references and contexts used to generate the answer.

Because `paperqa` splits the documents into chunks, each chunk is a valid reference. You can see that it also references the page where the context was found.

```python
print("3. References cited:")
print(f"{response.session.references}\n")
```

Lastly, `PQASession.session.contexts` contains the contexts used to generate the answer. Each context has a score, which is the similarity between the question and the context. `Paperqa` uses this score to choose what contexts is more relevant to answer the question.

```python
print("4. Contexts used to generate the answer:")
print(
    "These are the relevant text passages that were retrieved and used to formulate the answer:"
)
for i, ctx in enumerate(response.session.contexts, 1):
    print(f"\nContext {i}:")
    print(f"Source: {ctx.text.name}")
    print(f"Content: {ctx.context}")
    print(f"Score: {ctx.score}")
```


# Where to get papers

## OpenReview

You can use papers from <https://openreview.net/> as your database! Here's a helper that fetches a list of all papers from a selected conference (like ICLR, ICML, NeurIPS), queries this list to find relevant papers using LLM, and downloads those relevant papers to a local directory which can be used with paper-qa on the next step. Install `openreview-py` with

```bash
pip install paper-qa[openreview]
```

and get your username and password from the website. You can put them into `.env` file under `OPENREVIEW_USERNAME` and `OPENREVIEW_PASSWORD` variables, or pass them in the code directly.

```python
from paperqa import Settings
from paperqa.contrib.openreview_paper_helper import OpenReviewPaperHelper

# these settings require gemini api key you can get from https://aistudio.google.com/
# import os; os.environ["GEMINI_API_KEY"] = os.getenv("GEMINI_API_KEY")
# 1Mil context window helps to suggest papers. These settings are not required, but useful for an initial setup.
settings = Settings.from_name("openreview")
helper = OpenReviewPaperHelper(settings, venue_id="ICLR.cc/2025/Conference")
# if you don't know venue_id you can find it via
# helper.get_venues()

# Now we can query LLM to select relevant papers and download PDFs
question = "What is the progress on brain activity research?"

submissions = helper.fetch_relevant_papers(question)

# There's also a function that saves tokens by using openreview metadata for citations
docs = await helper.aadd_docs(submissions)

# Now you can continue asking like in the [main tutorial](../../README.md)
session = await docs.aquery(question, settings=settings)
print(session.answer)
```

## Zotero

*It's been a while since we've tested this - so let us know if it runs into issues!*

If you use [Zotero](https://www.zotero.org/) to organize your personal bibliography, you can use the `paperqa.contrib.ZoteroDB` to query papers from your library, which relies on [pyzotero](https://github.com/urschrei/pyzotero).

Install `pyzotero` via the `zotero` extra for this feature:

```bash
pip install paper-qa[zotero]
```

First, note that PaperQA2 parses the PDFs of papers to store in the database, so all relevant papers should have PDFs stored inside your database. You can get Zotero to automatically do this by highlighting the references you wish to retrieve, right clicking, and selecting *"Find Available PDFs"*. You can also manually drag-and-drop PDFs onto each reference.

To download papers, you need to get an API key for your account.

1. Get your library ID, and set it as the environment variable `ZOTERO_USER_ID`.
   * For personal libraries, this ID is given [here](https://www.zotero.org/settings/security#applications) at the part "*Your userID for use in API calls is XXXXXX*".
   * For group libraries, go to your group page `https://www.zotero.org/groups/groupname`, and hover over the settings link. The ID is the integer after /groups/. (*h/t pyzotero!*)
2. Create a new API key [here](https://www.zotero.org/settings/keys/new) and set it as the environment variable `ZOTERO_API_KEY`.
   * The key will need read access to the library.

With this, we can download papers from our library and add them to PaperQA2:

```python
from paperqa import Docs
from paperqa.contrib import ZoteroDB

docs = Docs()
zotero = ZoteroDB(library_type="user")  # "group" if group library

for item in zotero.iterate(limit=20):
    if item.num_pages > 30:
        continue  # skip long papers
    await docs.aadd(item.pdf, docname=item.key)
```

which will download the first 20 papers in your Zotero database and add them to the `Docs` object.

We can also do specific queries of our Zotero library and iterate over the results:

```python
for item in zotero.iterate(
    q="large language models",
    qmode="everything",
    sort="date",
    direction="desc",
    limit=100,
):
    print("Adding", item.title)
    await docs.aadd(item.pdf, docname=item.key)
```

You can read more about the search syntax by typing `zotero.iterate?` in IPython.

## Paper Scraper

If you want to search for papers outside of your own collection, I've found an unrelated project called [paper-scraper](https://github.com/blackadad/paper-scraper) that looks like it might help. But beware, this project looks like it uses some scraping tools that may violate publisher's rights or be in a gray area of legality.

```python
from paperqa import Docs

keyword_search = "bispecific antibody manufacture"
papers = paperscraper.search_papers(keyword_search)
docs = Docs()
for path, data in papers.items():
    try:
        await docs.aadd(path)
    except ValueError as e:
        # sometimes this happens if PDFs aren't downloaded or readable
        print("Could not read", path, e)
session = await docs.aquery(
    "What manufacturing challenges are unique to bispecific antibodies?"
)
print(session)
```


# packages


# paper-qa-pypdf

[![GitHub](https://img.shields.io/badge/GitHub-black?logo=github\&logoColor=white)](https://github.com/Future-House/paper-qa/tree/main/packages/paper-qa-docling) [![PyPI version](https://badge.fury.io/py/paper-qa-docling.svg)](https://badge.fury.io/py/paper-qa-docling) [![tests](https://github.com/Future-House/paper-qa/actions/workflows/tests.yml/badge.svg)](https://github.com/Future-House/paper-qa) ![License](https://img.shields.io/badge/License-Apache_2.0-blue.svg) ![PyPI Python Versions](https://img.shields.io/pypi/pyversions/paper-qa-docling)

PDF reading code backed by [Docling](https://github.com/docling-project/docling).

![docling logo](https://raw.githubusercontent.com/docling-project/docling/refs/heads/main/docs/assets/logo.png)

If you use this reader library in your projects, Docling requests you consider citing them: <https://github.com/docling-project/docling-parse/tree/main?tab=readme-ov-file#references>


# paper-qa-nemotron

[![GitHub](https://img.shields.io/badge/GitHub-black?logo=github\&logoColor=white)](https://github.com/Future-House/paper-qa/tree/main/packages/paper-qa-nemotron) [![tests](https://github.com/Future-House/paper-qa/actions/workflows/tests.yml/badge.svg)](https://github.com/Future-House/paper-qa) ![License](https://img.shields.io/badge/License-Apache_2.0-blue.svg) ![PyPI Python Versions](https://img.shields.io/pypi/pyversions/paper-qa-nemotron)

PDF reading code backed by [Nvidia's nemotron-parse VLM](https://build.nvidia.com/nvidia/nemotron-parse).

For more info on nemotron-parse, check out:

* Technical blog: <https://developer.nvidia.com/blog/turn-complex-documents-into-usable-data-with-vlm-nvidia-nemotron-parse-1-1/>
* Hugging Face weights: <https://huggingface.co/nvidia/NVIDIA-Nemotron-Parse-v1.1>
* NIM and model card: <https://build.nvidia.com/nvidia/nemotron-parse>
  * Support matrix: <https://docs.nvidia.com/nim/vision-language-models/1.5.0/support-matrix.html#nemotron-parse>
* API docs: <https://docs.nvidia.com/nim/vision-language-models/1.5.0/examples/nemotron-parse/overview.html#nemotron-parse-overview>
* Cookbook: <https://github.com/NVIDIA-NeMo/Nemotron/blob/main/usage-cookbook/Nemotron-Parse-v1.1/build_general_usage_cookbook.ipynb>
* NGC catalog: <https://catalog.ngc.nvidia.com/orgs/nim/teams/nvidia/containers/nemotron-parse>
* AWS Marketplace: <https://aws.amazon.com/marketplace/pp/prodview-ny2ngku2i4ge6>

## Installation

```bash
pip install paper-qa[nemotron]
# Or
pip install paper-qa-nemotron
```

If you want to prompt nemotron-parse hosted on AWS SageMaker:

```bash
pip install paper-qa-nemotron[sagemaker]
```

## Getting Started

To use nemotron-parse via the Nvidia API, set the `NVIDIA_API_KEY` environment variable.

Then to directly access the reader:

```python
from paperqa.types import ParsedText
from paperqa_nemotron import parse_pdf_to_pages

async def main(pdf_path) -> ParsedText:
    return await parse_pdf_to_pages(pdf_path)
```

Or use the reader within PaperQA:

```python
from paperqa import Docs, PQASession, Settings

from paperqa_nemotron import parse_pdf_to_pages


async def main(pdf_path, question: str | PQASession) -> PQASession:
    settings = Settings(parsing={"parse_pdf": parse_pdf_to_pages})
    docs = Docs()
    await docs.aadd(pdf_path, settings=settings)
    return await docs.aquery(question, settings=settings)
```


# paper-qa-pymupdf

[![GitHub](https://img.shields.io/badge/GitHub-black?logo=github\&logoColor=white)](https://github.com/Future-House/paper-qa/tree/main/packages/paper-qa-pymupdf) [![PyPI version](https://badge.fury.io/py/paper-qa-pymupdf.svg)](https://badge.fury.io/py/paper-qa-pymupdf) [![tests](https://github.com/Future-House/paper-qa/actions/workflows/tests.yml/badge.svg)](https://github.com/Future-House/paper-qa) ![License](https://img.shields.io/badge/license-AGPLv3-blue.svg) ![PyPI Python Versions](https://img.shields.io/pypi/pyversions/paper-qa-pymupdf)

PDF reading code backed by [PyMuPDF](https://github.com/pymupdf/PyMuPDF).


# paper-qa-pypdf

[![GitHub](https://img.shields.io/badge/GitHub-black?logo=github\&logoColor=white)](https://github.com/Future-House/paper-qa/tree/main/packages/paper-qa-pypdf) [![PyPI version](https://badge.fury.io/py/paper-qa-pypdf.svg)](https://badge.fury.io/py/paper-qa-pypdf) [![tests](https://github.com/Future-House/paper-qa/actions/workflows/tests.yml/badge.svg)](https://github.com/Future-House/paper-qa) ![License](https://img.shields.io/badge/License-Apache_2.0-blue.svg) ![PyPI Python Versions](https://img.shields.io/pypi/pyversions/paper-qa-pypdf)

PDF reading code backed by [PyPDF](https://github.com/py-pdf/pypdf).

To also parse images or take full-page screenshots, use the `media` extra: `pip install paper-qa-pypdf[media]`. This is backed by [pypdfium2](https://github.com/pypdfium2-team/pypdfium2).

From there, to also support grouping images into figures or parsing tables, use the `enhanced` extra: `pip install paper-qa-pypdf[enhanced]`. This is backed by [pdfplumber](https://github.com/jsvine/pdfplumber).


# Aviary

[![GitHub](https://img.shields.io/badge/github-%23121011.svg?style=for-the-badge\&logo=github\&logoColor=white)](https://github.com/Future-House/aviary) [![Project Status: Active](https://www.repostatus.org/badges/latest/active.svg)](https://www.repostatus.org/#active) ![License](https://img.shields.io/badge/License-Apache_2.0-blue.svg) [![Docs](https://assets.readthedocs.org/static/projects/badges/passing-flat.svg)](https://aviary.bio/) [![PyPI version](https://badge.fury.io/py/fhaviary.svg)](https://badge.fury.io/py/fhaviary) [![tests](https://github.com/Future-House/aviary/actions/workflows/tests.yml/badge.svg)](https://github.com/Future-House/aviary) [![CodeFactor](https://www.codefactor.io/repository/github/future-house/aviary/badge/main)](https://www.codefactor.io/repository/github/future-house/aviary/) [![Code style: black](https://img.shields.io/badge/code%20style-black-000000.svg)](https://github.com/psf/black) [![python](https://img.shields.io/badge/python-3.11%20%7C%203.12%20%7C%203.13-blue?style=flat\&logo=python\&logoColor=white)](https://www.python.org)

[![Crows in a gym](/files/PbEqj129hvvnBQPsFiNf)](https://arxiv.org/abs/2412.21154)

**Aviary** is a gymnasium for defining custom language agent RL environments. The library features pre-existing environments on math , general knowledge , biological sequences , scientific literature search , and protein stability. Aviary is designed to work in tandem with its sister library LDP (<https://github.com/Future-House/ldp>) which enables the user to define custom language agents as Language Decision Processes. See the following [tutorial](https://github.com/Future-House/aviary/blob/main/tutorials/Building%20a%20GSM8k%20Environment%20in%20Aviary.ipynb) for an example of how to run an LDP language agent on an Aviary environment.

[Overview](#overview) | [Getting Started](#getting-started) | [Documentation](https://aviary.bio/) | [Paper](https://arxiv.org/abs/2412.21154)

## What's New?

* We have a new environment to run Jupyter notebooks at [packages/notebook](/aviary/packages/notebook).

## Overview

[![Aviary and LDP overview from paper](/files/znycYw7yra1cs6VdwMm0)](https://arxiv.org/abs/2412.21154)

A pictorial overview of the five implemented Aviary environments and the language decision process framework.

## Getting Started

To install aviary (note `fh` stands for FutureHouse):

```bash
pip install fhaviary
```

To install aviary together with the incumbent environments:

```bash
pip install 'fhaviary[gsm8k,hotpotqa,labbench,lfrqa,notebook]'
```

To run the tutorial notebooks:

```bash
pip install "fhaviary[dev]"
```

### Developer Installation

For local development, please see [CONTRIBUTING.md](/aviary/contributing).

## Tutorial Notebooks

1. [Building a Custom Environment in Aviary](https://github.com/Future-House/aviary/blob/main/tutorials/Building%20a%20Custom%20Environment%20in%20Aviary.ipynb)
2. [Building a GSM8K Environment in Aviary](https://github.com/Future-House/aviary/blob/main/tutorials/Building%20a%20GSM8k%20Environment%20in%20Aviary.ipynb)
3. [Creating Language Agents to Interact with Aviary Environments](https://github.com/Future-House/ldp/blob/main/tutorials/creating_a_language_agent.ipynb)
4. [Evaluate a Llama Agent on GSM8K](https://github.com/Future-House/ldp/blob/main/tutorials/evaluating_a_llama_agent.ipynb)

## Defining a Custom Environment

The example below walks through defining a custom environment in Aviary. We define a simple environment where an agent takes actions to modify a counter. The example is also featured in the following [notebook](https://fastapi.tiangolo.com/advanced/path-operation-advanced-configuration/#advanced-description-from-docstring).

```python
from collections import namedtuple
from aviary.core import Environment, Message, ToolRequestMessage, Tool

# State in this example is simply a counter
CounterEnvState = namedtuple("CounterEnvState", ["count"])


class CounterEnv(Environment[CounterEnvState]):
    """A simple environment that allows an agent to modify a counter."""

    async def reset(self):
        """Initialize the environment with a counter set to 0. Goal is to count to 10"""
        self.state = CounterEnvState(count=0)

        # Target count
        self.target = 10

        # Create tools allowing the agent to increment and decrement counter
        self.tools = [
            Tool.from_function(self.incr),
            Tool.from_function(self.decr),
        ]

        # Return an observation message with the counter and available tools
        return [Message(content=f"Count to 10. counter={self.state.count}")], self.tools

    async def step(self, action: ToolRequestMessage):
        """Executes the tool call requested by the agent."""
        obs = await self.exec_tool_calls(action)

        # The reward is the square of the current count
        reward = int(self.state.count == self.target)

        # Returns observations, reward, done, truncated
        return obs, reward, reward == 1, False

    def incr(self):
        """Increment the counter."""
        self.state.count += 1
        return f"counter={self.state.count}"

    def decr(self):
        """Decrement the counter."""
        self.state.count -= 1
        return f"counter={self.state.count}"
```

## Evaluating an Agent on the Environment

Following the definition of our custom environment, we can now evaluate a language agent on the environment using Aviary's sister library LDP (<https://github.com/Future-House/ldp>).

```python
from ldp.agent import Agent
from ldp.graph import LLMCallOp
from ldp.alg import RolloutManager


class AgentState:
    """A container for maintaining agent state across interactions."""

    def __init__(self, messages, tools):
        self.messages = messages
        self.tools = tools


class SimpleAgent(Agent):
    def __init__(self, **kwargs):
        self._llm_call_op = LLMCallOp(**kwargs)

    async def init_state(self, tools):
        return AgentState([], tools)

    async def get_asv(self, agent_state, obs):
        """Take an action, observe new state, return value"""
        action = await self._llm_call_op(
            config={"name": "gpt-4o", "temperature": 0.1},
            msgs=agent_state.messages + obs,
            tools=agent_state.tools,
        )
        new_state = AgentState(
            messages=agent_state.messages + obs + [action],
            tools=agent_state.tools,
        )
        # Return action, state, value
        return action, new_state, 0.0


# Create a simple agent and perform rollouts on the environment

# Endpoint can be model identifier e.g. "claude-3-opus" depending on service
agent = SimpleAgent(config={"model": "my_llm_endpoint"})

runner = RolloutManager(agent=agent)

trajectories = await runner.sample_trajectories(
    environment_factory=CounterEnv,
    batch_size=2,
)
```

Below we expand on some of the core components of the Aviary library together with more advanced usage examples.

### Environment

An environment should have two methods, `env.reset` and `env.step`:

```py
obs_msgs, tools = await env.reset()
new_obs_msgs, reward, done, truncated = await env.step(action_msg)
```

Communication is achieved through messages.

The `action_msg` is an instance of `ToolRequestMessage` which comprises one or more calls to the `tools` returned by `env.reset` method.

The `obs_msgs` are either general obseravation messages or instances of `ToolResponseMessage` returned from the environment. while `reward` is a scalar value, and `done` and `truncated` are Boolean values.

We explain the message formalism in further detail below.

### Messages

Communication between the agent and environment is achieved via messages. We follow the [OpenAI](https://platform.openai.com/docs/api-reference/messages/createMessage) standard. Messages have two attributes:

```py
msg = Message(content="Hello, world!", role="assistant")
```

The `content` attribute can be a string but can also comprise objects such as [images](https://platform.openai.com/docs/guides/vision?lang=node#uploading-base64-encoded-images). For example, the `create_message` method can be used to create a message with images:

```py
from PIL import Image
import numpy as np

img = Image.open("your_image.jpg")
img_array = np.array(img)

msg = Message.create_message(role="user", text="Hello, world!", images=[img_array])
```

In this case, `content` will be a list of dictionaries with the keys `text` and `image_url`.

```py
{
    {"type": "text", "text": "Hello World!"},
    {"text": "image_url", "image_url": "data:image/png;base64,{base64_image}"},
}
```

The role, see the table below. You can change around roles as desired, except for `tool` which has a special meaning in aviary.

| Role      | Host                                             | Example(s)                                                       |
| --------- | ------------------------------------------------ | ---------------------------------------------------------------- |
| assistant | Agent                                            | An agent's tool selection message                                |
| system    | Agent system prompt                              | "You are an agent."                                              |
| user      | Environment system prompt or emitted observation | HotPotQA problem to solve, or details of an internal env failure |
| tool      | Result of a tool run in the environment          | The output of the calculator tool for a GSM8K question           |

The `Message` class is extended in `ToolRequestMessage` and `ToolResponseMessage` to include the relevant tool name and arguments.

### Subclassing Environments

If you need more control over Environments and tools, you may wish to subclass `Environment`. We illustrate this with an example environment in which an agent is tasked to write a story.

We subclass `Environment` and define a `state`. The `state` consists of all variables that change per step that we wish to bundle together. It will be accessible in tools, so you can use `state` to store information you want to persist between steps and tool calls.

```py
from pydantic import BaseModel
from aviary.core import Environment


class ExampleState(BaseModel):
    reward: float = 0
    done: bool = False


class ExampleEnv(Environment[ExampleState]):
    state: ExampleState
```

We do not have other variables aside from `state` for this environment, although we could also have variables like configuration, a name, tasks, etc. attached to it.

### Defining Tools

We will define a single tool that prints a story. Tools may optionally take a final argument `state` which is the environment state. This argument will not be exposed to the agent as a parameter but will be injected by the environment (if part of the function signature).

```py
def print_story(story: str, state: ExampleState):
    """Print a story.

    Args:
        story: Story to print.
        state: Environment state (hidden from agent).
    """
    print(story)
    state.reward = 1
    state.done = True
```

The tool is built from the following parts of the function: its name, its argument's names, the arguments types, and the docstring. The docstring is parsed to obtain a description of the function and its arguments, so be sure to match the syntax carefully.

Environment episode completion is indicated by setting `state.done = True`. This example terminates immediately - other termination conditions are also possible.

It is also possible make the function `async` - the environment will account for that when the tool is called.

### Advanced Tool Descriptions

Aviary also supports more sophisticated signatures:

* Multiline docstrings
* Non-primitive type hints (e.g. type unions)
* Default values
* Exclusion of info below `\f` (see below)

If you have summary-level information that belongs in the docstring, but you don't want it to be part of the `Tool.info.description`, add a `r` prefix to the docstring and inject `\f` before the summary information to exclude. This convention was created by FastAPI ([docs](https://fastapi.tiangolo.com/advanced/path-operation-advanced-configuration/#advanced-description-from-docstring)).

```python
def print_story(story: str | bytes, state: ExampleState):
    r"""Print a story.

    Extra information that is part of the tool description.

    \f

    This sentence is excluded because it's an implementation detail.

    Args:
        story: Story to print, either as a string or bytes.
        state: Environment state.
    """
    print(story)
    state.reward = 1
    state.done = True
```

### The Environment `reset` Method

Next we define the `reset` function which initializes the tools and returns one or more initial observations as well as the tools. The `reset` function is `async` to allow for database interactions or HTTP requests.

```py
from aviary.core import Message, Tool


async def reset(self):
    self.tools = [Tool.from_function(ExampleEnv.print_story)]
    start = Message(content="Write a 5 word story and call print")
    return [start], self.tools
```

### The Environment `step` Method

Next we define the `step` function which takes an action and returns the next observation, reward, done, and whether the episode was truncated.

```py
from aviary.core import Message


async def step(self, action: Message):
    msgs = await self.exec_tool_calls(action, state=self.state)
    return msgs, self.state.reward, self.state.done, False
```

You will probably often use this specific syntax for calling the tools - calling `exec_tool_calls` with the action.

### Environment `export_frame` Method

Optionally, we can define a function to export a snapshot of the environment and its state for visualization or debugging purposes.

```py
from aviary.core import Frame


def export_frame(self):
    return Frame(
        state={"done": self.state.done, "reward": self.state.reward},
        info={"tool_names": [t.info.name for t in self.tools]},
    )
```

### Viewing Environment Tools

If an environment can be instantiated without anything other than the task (i.e., it implements `from_task`), you can start a server to view its tools:

```sh
pip install fhaviary[server]
aviary tools [env name]
```

This will start a server that allows you to view the tools and call them, viewing the descriptions/types and output that an agent would see when using the tools.

### Incumbent Environments

Below we list some pre-existing environments implemented in Aviary:

| Environment | PyPI                                                           | Extra                | README                                                |
| ----------- | -------------------------------------------------------------- | -------------------- | ----------------------------------------------------- |
| GSM8k       | [`aviary.gsm8k`](https://pypi.org/project/aviary.gsm8k/)       | `fhaviary[gsm8k]`    | [`README.md`](/aviary/packages/gsm8k#installation)    |
| HotPotQA    | [`aviary.hotpotqa`](https://pypi.org/project/aviary.hotpotqa/) | `fhaviary[hotpotqa]` | [`README.md`](/aviary/packages/hotpotqa#installation) |
| LAB-Bench   | [`aviary.labbench`](https://pypi.org/project/aviary.labbench/) | `fhaviary[labbench]` | [`README.md`](/aviary/packages/labbench#installation) |
| LFRQA       | [`aviary.lfrqa`](https://pypi.org/project/aviary.lfrqa/)       | `fhaviary[lfrqa]`    | [`README.md`](/aviary/packages/lfrqa#installation)    |
| Notebook    | [`aviary.notebook`](https://pypi.org/project/aviary.notebook/) | `fhaviary[notebook]` | [`README.md`](/aviary/packages/notebook#installation) |
| LitQA       | [`aviary.litqa`](https://pypi.org/project/aviary.litqa/)       | Moved to `labbench`  | Moved to `labbench`                                   |

### Task Datasets

Included with some environments are collections of problems that define training or evaluation datasets. We refer to these as `TaskDataset`s, e.g. for the `HotpotQADataset` subclass of `TaskDataset`:

```py
from aviary.envs.hotpotqa import HotPotQADataset

dataset = HotPotQADataset(split="dev")
```

### Functional Environments

An alternative way to create an environment is using the functional interface, which uses functions and decorators to define environments. Let's define an environment that requires an agent to write a story about a particular topic by implementing its `start` function:

```python
from aviary.core import fenv


@fenv.start()
def my_env(topic):
    # return the first observation and starting environment state
    # (empty in this case)
    return f"Write a story about {topic}", {}
```

The `start` decorator begins the definition of an environment.

The function, `my_env`, takes an arbitrary input and returns a tuple containing the first observation and any information you wish to store about the environment state (used to persist/share information between tools).

The state will always have an optional `reward` and a Boolean `done` that indicate if the environment episode is complete. Next we define some tools:

```python
@my_env.tool()
def multiply(x: float, y: float) -> float:
    """Multiply two numbers."""
    return x * y


@my_env.tool()
def print_story(story: str | bytes, state) -> None:
    """Print a story to the user and complete episode."""
    print(story)
    state.reward = 1
    state.done = True
```

The tools will be converted into objects visible for LLMs using the type hints and the variable descriptions. Thus, the type hinting can be valuable for an agent that uses it correctly. The docstrings are also passed to the LLM and is the primary means (along with the function name) for communicating the intended tool usage.

You can access the `state` variable in tools, which will have any fields you passed in the return tuple of `start()`. For example, if you returned `{'foo': 'bar'}`, then you could access `state.foo` in the tools.

You may stop an environment or set a reward via the `state` variable as shown in the second `print_story` tool. If the reward is not set, it is treated as zero. Next we illustrate how to use our environment:

```python
env = my_env(topic="foo")
obs, tools = await env.reset()
```

## Citing Aviary

If Aviary is useful for your work please consider citing the following paper:

```bibtex
@article{Narayanan_Aviary_training_language_2024,
  title   = {{Aviary: training language agents on challenging scientific tasks}},
  author  = {
    Narayanan, Siddharth and Braza, James D. and Griffiths, Ryan-Rhys and
    Ponnapati, Manvitha and Bou, Albert and Laurent, Jon and Kabeli, Ori and
    Wellawatte, Geemi and Cox, Sam and Rodriques, Samuel G. and White, Andrew
    D.
  },
  year    = 2024,
  month   = dec,
  journal = {preprint},
  doi     = {10.48550/arXiv.2412.21154},
  url     = {https://arxiv.org/abs/2412.21154}
}
```

## References


# Contributing to aviary

## Repo Structure

aviary is a monorepo using [`uv`'s workspace layout](https://docs.astral.sh/uv/concepts/workspaces/#workspace-layouts).

## Installation

1. Git clone this repo
2. Install the project manager `uv`: <https://docs.astral.sh/uv/getting-started/installation/>
3. Run `uv sync`

This will editably install the full monorepo in your local environment.

## Testing

To run tests, please just run `pytest` in the repo root.

Note you will need OpenAI and Anthropic API keys configured.


# packages


# aviary.gsm8k

GSM8k environment where agents solve math word problems from the GSM8k dataset using a calculator tool.

## Citation

The citation for GSM8k is given below:

```bibtex
@article{gsm8k-paper,
  title   = {Training verifiers to solve math word problems},
  author  = {
    Cobbe, Karl and Kosaraju, Vineet and Bavarian, Mohammad and Chen, Mark and
    Jun, Heewoo and Kaiser, Lukasz and Plappert, Matthias and Tworek, Jerry and
    Hilton, Jacob and Nakano, Reiichiro and others
  },
  year    = 2021,
  journal = {arXiv preprint arXiv:2110.14168}
}
```

## References

\[1] Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R. and Hesse, C., 2021. [Training verifiers to solve math word problems](https://arxiv.org/abs/2110.14168). arXiv preprint arXiv:2110.14168.

## Installation

To install the GSM8k environment, run the following command:

```bash
pip install 'fhaviary[gsm8k]'
```


# aviary.hotpotqa

The HotPotQA environment asks agents to perform multi-hop question answering on the HotPotQA dataset \[1].

## References

\[1] Yang et al. [HotpotQA: A Dataset for Diverse,\
Explainable Multi-Hop Question Answering](https://aclanthology.org/D18-1259/). EMNLP, 2018.

## Installation

To install the HotPotQA environment, run the following command:

```bash
pip install 'fhaviary[hotpotqa]'
```


# aviary.labbench

LAB-Bench environments implemented with aviary, allowing agents to perform question answering on scientific tasks.

## Installation

To install the LAB-Bench environment, run:

```bash
pip install 'fhaviary[labbench]'
```

## Usage

In [`labbench/env.py`](https://github.com/Future-House/aviary/blob/main/packages/labbench/src/aviary/envs/labbench/env.py), you will find:

* `GradablePaperQAEnvironment`: an PaperQA-backed environment that can grade answers given an evaluation function.
* `ImageQAEnvironment`: an `GradablePaperQAEnvironment` subclass for QA where image(s) are pre-added.

And in [`labbench/task.py`](https://github.com/Future-House/aviary/blob/main/packages/labbench/src/aviary/envs/labbench/task.py), you will find:

* `TextQATaskDataset`: a task dataset designed to pull down FigQA, LitQA2, or TableQA from Hugging Face, and create one `GradablePaperQAEnvironment` per question.
* `ImageQATaskDataset`: a task dataset that pairs with `ImageQAEnvironment` for FigQA or TableQA.

Here is an example of how to use them:

```python
import os

from ldp.agent import SimpleAgent
from ldp.alg import Evaluator, EvaluatorConfig, MeanMetricsCallback
from paperqa import Settings

from aviary.env import TaskDataset


async def evaluate(folder_of_litqa_v2_papers: str | os.PathLike) -> None:
    settings = Settings(paper_directory=folder_of_litqa_v2_papers)
    dataset = TaskDataset.from_name("litqa2", settings=settings)
    metrics_callback = MeanMetricsCallback(eval_dataset=dataset)

    evaluator = Evaluator(
        config=EvaluatorConfig(batch_size=3),
        agent=SimpleAgent(),
        dataset=dataset,
        callbacks=[metrics_callback],
    )
    await evaluator.evaluate()
    print(metrics_callback.eval_means)
```

### Image Question-Answer

This is an environment/dataset for giving PaperQA a `Docs` object with the image(s) for one LAB-Bench question. It's designed to be a comparison with zero-shotting the question to a LLM, but instead of a singular prompt the image is put through the PaperQA agent loop.

```python
from typing import cast

import litellm
import pytest
from ldp.agent import Agent
from ldp.alg import (
    Evaluator,
    EvaluatorConfig,
    MeanMetricsCallback,
    StoreTrajectoriesCallback,
)
from paperqa.settings import AgentSettings, IndexSettings

from aviary.envs.labbench import (
    ImageQAEnvironment,
    ImageQATaskDataset,
    LABBenchDatasets,
)


@pytest.mark.asyncio
async def test_image_qa(tmp_path) -> None:
    litellm.num_retries = 8  # Mitigate connection-related failures
    settings = ImageQAEnvironment.make_base_settings()
    settings.agent = AgentSettings(
        agent_type="ldp.agent.SimpleAgent",
        index=IndexSettings(paper_directory=tmp_path),
        # TODO: add image support for paper_search
        tool_names={"gather_evidence", "gen_answer", "complete", "reset"},
        agent_evidence_n=3,  # Bumped up to collect several perspectives
    )
    dataset = ImageQATaskDataset(dataset=LABBenchDatasets.TABLE_QA, settings=settings)
    t_cb = StoreTrajectoriesCallback()
    m_cb = MeanMetricsCallback(eval_dataset=dataset, track_tool_usage=True)
    evaluator = Evaluator(
        config=EvaluatorConfig(
            batch_size=256,  # Use batch size greater than FigQA size and TableQA size
            max_rollout_steps=18,  # Match aviary paper's PaperQA setting
        ),
        agent=cast(Agent, await settings.make_ldp_agent(settings.agent.agent_type)),
        dataset=dataset,
        callbacks=[t_cb, m_cb],
    )
    await evaluator.evaluate()
    print(m_cb.eval_means)
```

## References

\[1] Skarlinski et al. [Language agents achieve superhuman synthesis of scientific knowledge](https://arxiv.org/abs/2409.13740). ArXiv:2409.13740, 2024.

\[2] Laurent et al. [LAB-Bench: Measuring Capabilities of Language Models for Biology Research](https://arxiv.org/abs/2407.10362). ArXiv:2407.10362, 2024.


# aviary.lfrqa

An environment designed to utilize PaperQA for answering questions from the LFRQATaskDataset Long-form RobustQA (LFRQA) is a human-annotated dataset introduced in the RAG-QA-Arena, featuring over 1400 questions from various categories, including science.

## Installation

To install the LFRQA environment, run:

```bash
pip install 'fhaviary[lfrqa]'
```

## Usage

Refer to [this tutorial](https://github.com/Future-House/paper-qa/blob/main/docs/tutorials/running_on_lfrqa.md) for instructions on how to run the environment.

## References

\[1] RAG-QA Arena (<https://arxiv.org/pdf/2407.13998>)


# aviary.notebook

A Jupyter notebook environment.

## Installation

To install the notebook environment, run the following command:

```bash
pip install 'fhaviary[notebook]'
```

To allow the environment to run notebooks in containerized sandboxes (recommended), first build the default image:

```bash
cd docker/
docker build -t aviary-notebook-env -f Dockerfile.pinned .
```

And second, set the environment variable `NB_ENVIRONMENT_USE_DOCKER=true`. You may use your own Docker image with a different name, in which case you must override the environment variable `NB_ENVIRONMENT_DOCKER_IMAGE`.


# Language Decision Processes (LDP)

[![GitHub](https://img.shields.io/badge/github-%23121011.svg?style=for-the-badge\&logo=github\&logoColor=white)](https://github.com/Future-House/ldp) [![Project Status: Active](https://www.repostatus.org/badges/latest/active.svg)](https://www.repostatus.org/#active) ![License](https://img.shields.io/badge/License-Apache_2.0-blue.svg) [![Docs](https://assets.readthedocs.org/static/projects/badges/passing-flat.svg)](https://futurehouse.gitbook.io/futurehouse-cookbook/ldp-language-decision-processes) [![PyPI version](https://badge.fury.io/py/ldp.svg)](https://badge.fury.io/py/ldp) [![tests](https://github.com/Future-House/ldp/actions/workflows/tests.yml/badge.svg)](https://github.com/Future-House/ldp) [![CodeFactor](https://www.codefactor.io/repository/github/future-house/ldp/badge)](https://www.codefactor.io/repository/github/future-house/ldp) [![Code style: black](https://img.shields.io/badge/code%20style-black-000000.svg)](https://github.com/psf/black) [![python](https://img.shields.io/badge/python-3.11%20%7C%203.12%20%7C%203.13-blue?style=flat\&logo=python\&logoColor=white)](https://www.python.org)

<p align="center"><a href="https://arxiv.org/abs/2412.21154"><img src="/files/sdk8ovSq0UKizHryZm7D" alt="row playing chess"></a></p>

**LDP** is a framework for enabling modular interchange of language agents, environments, and optimizers. A language decision process (LDP) is a partially-observable Markov decision process (POMDP) where actions and observations consist of natural language. The full definition from the Aviary paper is:

[![LDP definition from paper](/files/K725kxuTZS2Zt3uCYCTn)](https://arxiv.org/abs/2412.21154)

See the following [tutorial](https://github.com/Future-House/ldp/blob/main/tutorials/creating_a_language_agent.ipynb) for an example of how to run an LDP agent.

[Overview](#overview) | [Getting Started](#getting-started) | [Documentation](https://futurehouse.gitbook.io/futurehouse-cookbook/ldp-language-decision-processes) | [Paper](https://arxiv.org/abs/2412.21154)

## What's New?

* Check out our new [Tutorial](https://github.com/Future-House/ldp/blob/main/tutorials/creating_a_language_agent.ipynb) notebook on running an LDP agent in an Aviary environment!
* The Aviary paper has been posted to [arXiv](https://arxiv.org/abs/2412.21154)! Further updates forthcoming!

## Overview

[![Aviary and LDP overview from paper](/files/WpzR0E2ADGqWrQJ4jRZ2)](https://arxiv.org/abs/2412.21154)

A pictorial overview of the language decision process (LDP) framework together with five implemented Aviary environments.

## Getting Started

To install `ldp`:

```bash
pip install -e .
```

To install `aviary` and the `nn` (neural network) module required for the tutorials:

```bash
pip install "ldp[nn]" "fhaviary[gsm8k]"
```

If you plan to export Graphviz visualizations, the `graphviz` library is required:

* Linux: `apt install graphviz`
* macOS: `brew install graphviz`

## Tutorial Notebooks

1. [Creating a Simple Language Agent](https://github.com/Future-House/ldp/blob/main/tutorials/creating_a_language_agent.ipynb)
2. [Evaluating a Llama Agent on GSM8K](https://github.com/Future-House/ldp/blob/main/tutorials/evaluating_a_llama_agent.ipynb)

## Running an Agent on an Aviary Environment

The minimal example below illustrates how to run a language agent on an Aviary environment (LDP's sister library for defining language agent environments - <https://github.com/Future-House/aviary>)

```py
from ldp.agent import SimpleAgent
from aviary.core import DummyEnv

env = DummyEnv()
agent = SimpleAgent()

obs, tools = await env.reset()
agent_state = await agent.init_state(tools=tools)

done = False
while not done:
    action, agent_state, _ = await agent.get_asv(agent_state, obs)
    obs, reward, done, truncated = await env.step(action.value)
```

Below we elaborate on the components of LDP.

## Agent

An agent is a language agent that interacts with an environment to accomplish a task. Agents may use tools (calls to external APIs e.g. Wolfram Alpha) in response to observations returned by the environment. Below we define LDP's `SimpleAgent` which relies on a single LLM call. The main bookkeeping involves appending messages received from the environment and passing tools.

```py
from ldp.agent import Agent
from ldp.graph import LLMCallOp


class AgentState:
    def __init__(self, messages, tools):
        self.messages = messages
        self.tools = tools


class SimpleAgent(Agent):
    def __init__(self, **kwargs):
        super().__init__(**kwargs)
        self.llm_call_op = LLMCallOp()

    async def init_state(self, tools):
        return AgentState([], tools)

    async def get_asv(self, agent_state, obs):
        action = await self.llm_call_op(
            config={"name": "gpt-4o", "temperature": 0.1},
            msgs=agent_state.messages + obs,
            tools=agent_state.tools,
        )
        new_state = AgentState(
            messages=agent_state.messages + obs + [action], tools=agent_state.tools
        )
        return action, new_state, 0.0
```

An agent has two methods:

```py
agent_state = await agent.init_state(tools=tools)
new_action, new_agent_state, value = await agent.get_asv(agent_state, obs)
```

* The `get_asv(agent_state, obs)` method chooses an action (`a`) conditioned on the observation messages returning the next agent state (`s`) and a value estimate (`v`).
* The first argument, `agent_state`, is an optional container for environment-specific objects such as e.g. documents for PaperQA or lookup results for HotpotQA,
* as well as more general objects such as memories which could include a list of previous actions and observations. `agent_state` may be set to `None` if memories are not being used.
* The second argument `obs` is not the complete list of all prior observations, but rather the returned value from `env.step`.
* The `value` is the agent's state/action value estimate used for reinforcment learning training. It may default to 0.

## A plain python agent

Want to just run python code? No problem - here's a minimal example of an Agent that is deterministic:

```py
from aviary.core import Message, Tool, ToolCall, ToolRequestMessage
from ldp.agent import Agent


class NoThinkAgent(Agent):
    async def init_state(self, tools):
        return None

    async def get_asv(self, tools, obs):
        tool_call = ToolCall.from_name("specific_tool_call", arg1="foo")
        action = ToolRequestMessage(tool_calls=[tool_call])
        return await Agent.wrap_action(action), None, 0.0
```

This agent has a state of `None`, just makes one specific tool call with `arg1="foo"`, and then converts that into an action. The only "magic" line of code is the `wrap_action`, which just converts the action constructed by plain python into a node in a compute graph - see more below.

## Stochastic Computation Graph (SCG)

For more advanced use-cases, LDP features a stochastic computation graph which enables differentiatiation with respect to agent parameters (including the weights of the LLM).

You should install the `scg` subpackage to work with it:

```bash
pip install ldp[scg]
```

The example computation graph below illustrates the functionality

```py
from ldp.graph import FxnOp, LLMCallOp, PromptOp, compute_graph

op_a = FxnOp(lambda x: 2 * x)

async with compute_graph():
    op_result = op_a(3)
```

The code cell above creates and executes a computation graph that doubles the input. The computation graph gradients and executions are saved in a context for later use, such as in training updates. For example:

```py
print(op_result.compute_grads())
```

A more complex example is given below for an agent that possesses memory.

```py
@compute_graph()
async def get_asv(self, agent_state, obs):
    # Update state with new observations
    next_state = agent_state.get_next_state(obs)

    # Retrieve relevant memories
    query = await self._query_factory_op(next_state.messages)
    memories = await self._memory_op(query, matches=self.num_memories)

    # Format memories and package messages
    formatted_memories = await self._format_memory_op(self.memory_prompt, memories)
    memory_prompt = await self._prompt_op(memories=formatted_memories)
    packaged_messages = await self._package_op(
        next_state.messages, memory_prompt=memory_prompt, use_memories=bool(memories)
    )

    # Make LLM call and update state
    config = await self._config_op()
    result = await self._llm_call_op(
        config, msgs=packaged_messages, tools=next_state.tools
    )
    next_state.messages.extend([result])

    return result, next_state, 0.0
```

We use differentiable ops to ensure there is an edge in the compute graph from the LLM result (action) to components such as memory retrieval as well as the query used to retrieve the memory.

Why use an SCG? Aside from the ability to take gradients, using the SCG enables tracking of all inputs/outputs to the ops and serialization/deserialization of the SCG such that it can be easily saved and loaded. Input/output tracking also makes it easier to perform fine-tuning or reinforcement learning on the underlying LLMs.

## Generic Support

The `Agent` (as well as classes in `agent.ops`) are [generics](https://en.wikipedia.org/wiki/Generic_programming), which means:

* `Agent` is designed to support arbitrary types
* Subclasses can precisely specify state types, making the code more readable

If you are new to Python generics (`typing.Generic`), please read about them in [Python `typing`](https://docs.python.org/3/library/typing.html#generics). Below is how to specify an agent with a custom state type.

```py
from dataclasses import dataclass, field
from datetime import datetime

from ldp.agents import Agent


@dataclass
class MyComplexState:
    vector: list[float]
    timestamp: datetime = field(default_factory=datetime.now)


class MyAgent(Agent[MyComplexState]):
    """Some agent who is now type checked to match the custom state."""
```

## References


# packages


# Language Model Interface (LMI)

[![GitHub](https://img.shields.io/badge/github-%23121011.svg?style=for-the-badge\&logo=github\&logoColor=white)](/ldp-language-decision-processes/packages/lmi) [![PyPI version](https://badge.fury.io/py/fhlmi.svg)](https://badge.fury.io/py/fhlmi) [![tests](https://github.com/Future-House/ldp/actions/workflows/tests.yml/badge.svg)](/ldp-language-decision-processes/packages/lmi) ![License](https://img.shields.io/badge/License-Apache_2.0-blue.svg) ![PyPI Python Versions](https://img.shields.io/pypi/pyversions/fhlmi)

A Python library for interacting with Large Language Models (LLMs) through a unified interface, hence the name Language Model Interface (LMI).

## Installation

```bash
pip install fhlmi
```

***

**Table of Contents**

* [Installation](#installation)
* [Quick start](#quick-start)
* [Documentation](#documentation)
  * [LLMs](#llms)
    * [LLMModel](#llmmodel)
    * [LiteLLMModel](#litellmmodel)
  * [Cost tracking](#cost-tracking)
  * [Rate limiting](#rate-limiting)
    * [Basic Usage](#basic-usage)
    * [Rate Limit Format](#rate-limit-format)
    * [Storage Options](#storage-options)
    * [Monitoring Rate Limits](#monitoring-rate-limits)
    * [Timeout Configuration](#timeout-configuration)
    * [Weight-based Rate Limiting](#weight-based-rate-limiting)
  * [Tool calling](#tool-calling)
  * [Vertex](#vertex)
  * [Embedding models](#embedding-models)
    * [LiteLLMEmbeddingModel](#litellmembeddingmodel)
    * [HybridEmbeddingModel](#hybridembeddingmodel)
    * [SentenceTransformerEmbeddingModel](#sentencetransformerembeddingmodel)

***

## Quick start

A simple example of how to use the library with default settings is shown below.

```python
from lmi import LiteLLMModel
from aviary.core import Message

llm = LiteLLMModel()

messages = [Message(content="What is the meaning of life?")]

result = await llm.call_single(messages)
# assert result.text == "42"
```

or, if you only have one user message, just:

```python
from lmi import LiteLLMModel

llm = LiteLLMModel()
result = await llm.call_single("What is the meaning of life?")
# assert result.text == "42"
```

## Documentation

### LLMs

An LLM is a class that inherits from `LLMModel` and implements the following methods:

* `async acompletion(messages: list[Message], **kwargs) -> list[LLMResult]`
* `async acompletion_iter(messages: list[Message], **kwargs) -> AsyncIterator[LLMResult]`

These methods are used by the base class `LLMModel` to implement the LLM interface. Because `LLMModel` is an abstract class, it doesn't depend on any specific LLM provider. All the connection with the provider is done in the subclasses using `acompletion` and `acompletion_iter` as interfaces.

Because these are the only methods that communicate with the chosen LLM provider, we use an abstraction [LLMResult](https://github.com/Future-House/ldp/blob/main/packages/lmi/src/lmi/types.py#L35) to hold the results of the LLM call.

#### LLMModel

An `LLMModel` implements `call`, which receives a list of `aviary` `Message`s and returns a list of `LLMResult`s. `LLMModel.call` can receive callbacks, tools, and output schemas to control its behavior, as better explained below. Because we support interacting with the LLMs using `Message` objects, we can use the modalities available in `aviary`, which currently include text and images. `lmi` supports these modalities but does not support other modalities yet. Adittionally, `LLMModel.call_single` can be used to return a single `LLMResult` completion.

#### LiteLLMModel

`LiteLLMModel` wraps `LiteLLM` API usage within our `LLMModel` interface. It receives a `name` parameter, which is the name of the model to use and a `config` parameter, which is a dictionary of configuration options for the model following the [LiteLLM configuration schema](https://docs.litellm.ai/docs/routing). Common parameters such as `temperature`, `max_token`, and `n` (the number of completions to return) can be passed as part of the `config` dictionary.

```python
import os
from lmi import LiteLLMModel

config = {
    "model_list": [
        {
            "model_name": "gpt-4o",
            "litellm_params": {
                "model": "gpt-4o",
                "api_key": os.getenv("OPENAI_API_KEY"),
                "frequency_penalty": 1.5,
                "top_p": 0.9,
                "max_tokens": 512,
                "temperature": 0.1,
                "n": 5,
            },
        }
    ]
}

llm = LiteLLMModel(name="gpt-4o", config=config)
```

`config` can also be used to pass common parameters directly for the model.

```python
from lmi import LiteLLMModel

config = {
    "name": "gpt-4o",
    "temperature": 0.1,
    "max_tokens": 512,
    "n": 5,
}

llm = LiteLLMModel(config=config)
```

### Cost tracking

Cost tracking is supported in two different ways:

1. Calls to the LLM return the token usage for each call in `LLMResult.prompt_count` and `LLMResult.completion_count`. Additionally, `LLMResult.cost` can be used to get a cost estimate for the call in USD.
2. A global cost tracker is maintained in `GLOBAL_COST_TRACKER` and can be enabled or disabled using `enable_cost_tracking()` and `cost_tracking_ctx()`.

### Rate limiting

Rate limiting helps regulate the usage of resources to various services and LLMs. The rate limiter supports both in-memory and Redis-based storage for cross-process rate limiting. Currently, `lmi` take into account the tokens used (Tokens per Minute (TPM)) and the requests handled (Requests per Minute (RPM)).

#### Basic Usage

Rate limits can be configured in two ways:

1. Through the LLM configuration:

   ```python
   from lmi import LiteLLMModel

   config = {
       "rate_limit": {
           "gpt-4": "100/minute",  # 100 tokens per minute
       },
       "request_limit": {
           "gpt-4": "5/minute",  # 5 requests per minute
       },
   }

   llm = LiteLLMModel(name="gpt-4", config=config)
   ```

   With `rate_limit` we rate limit only token consumption, and with `request_limit` we rate limit only request volume. You can configure both of them or only one of them as you need.
2. Through the global rate limiter configuration:

   ```python
   from lmi.rate_limiter import GLOBAL_LIMITER

   GLOBAL_LIMITER.rate_config[("client", "gpt-4")] = "100/minute"  # tokens per minute
   GLOBAL_LIMITER.rate_config[("client|request", "gpt-4")] = (
       "5/minute"  # requests per minute
   )
   ```

   With `client` we rate limit only token consumption, and with `client|request` we rate limit only request volume. You can configure both of them or only one of them as you need.

#### Rate Limit Format

Rate limits can be specified in two formats:

1. As a string: `"<count> [per|/] [n (optional)] <second|minute|hour|day|month|year>"`

   ```python
   "100/minute"  # 100 tokens per minute

   "5 per second"  # 5 tokens per second
   "1000/day"  # 1000 tokens per day
   ```
2. Using RateLimitItem classes:

   ```python
   from limits import RateLimitItemPerSecond, RateLimitItemPerMinute

   RateLimitItemPerSecond(30, 1)  # 30 tokens per second
   RateLimitItemPerMinute(1000, 1)  # 1000 tokens per minute
   ```

#### Storage Options

The rate limiter supports two storage backends:

1. In-memory storage (default when Redis is not configured):

   ```python
   from lmi.rate_limiter import GlobalRateLimiter

   limiter = GlobalRateLimiter(use_in_memory=True)
   ```
2. Redis storage (for cross-process rate limiting):

   ```python
   # Set REDIS_URL environment variable
   import os

   os.environ["REDIS_URL"] = "redis://localhost:6379"

   from lmi.rate_limiter import GlobalRateLimiter

   limiter = GlobalRateLimiter()  # Will automatically use Redis if REDIS_URL is set
   ```

   This `limiter` can be used in within the `LLMModel.check_rate_limit` method to check the rate limit before making a request, similarly to how it is done in the [`LiteLLMModel` class](https://github.com/Future-House/ldp/blob/18138af155bef7686d1eb2b486edbc02d62037eb/packages/lmi/src/lmi/llms.py).

#### Monitoring Rate Limits

You can monitor current rate limit status:

```python
from lmi.rate_limiter import GLOBAL_LIMITER
from lmi import LiteLLMModel
from aviary.core import Message

config = {
    "rate_limit": {
        "gpt-4": "100/minute",  # 100 tokens per minute
    },
    "request_limit": {
        "gpt-4": "5/minute",  # 5 requests per minute
    },
}

llm = LiteLLMModel(name="gpt-4", config=config)
results = await llm.call([Message(content="Hello, world!")])  # Consume some tokens

status = await GLOBAL_LIMITER.rate_limit_status()

# Example output:
{
    ("client|request", "gpt-4"): {  # the limit status for requests
        "period_start": 1234567890,
        "n_items_in_period": 1,
        "period_seconds": 60,
        "period_name": "minute",
        "period_cap": 5,
    },
    ("client", "gpt-4"): {  # the limit status for tokens
        "period_start": 1234567890,
        "n_items_in_period": 50,
        "period_seconds": 60,
        "period_name": "minute",
        "period_cap": 100,
    },
}
```

#### Timeout Configuration

The default timeout for rate limiting is 60 seconds, but can be configured:

```python
import os

os.environ["RATE_LIMITER_TIMEOUT"] = "30"  # 30 seconds timeout
```

#### Weight-based Rate Limiting

Rate limits can account for different weights (e.g., token counts for LLM requests):

```python
await GLOBAL_LIMITER.try_acquire(
    ("client", "gpt-4"),
    weight=token_count,  # Number of tokens in the request
    acquire_timeout=30.0,  # Optional timeout override
)
```

### Tool calling

LMI supports function calling through tools, which are functions that the LLM can invoke. Tools are passed to `LLMModel.call` or `LLMModel.call_single` as a list of [`Tool` objects from `aviary`](https://github.com/Future-House/aviary/blob/1a50b116fb317c3ef27b45ea628781eb53c0b7ae/src/aviary/tools/base.py#L334), along with an optional `tool_choice` parameter that controls how the LLM uses these tools.

The `tool_choice` parameter follows `OpenAI`'s definition. It can be:

| Tool Choice Value               | Constant                           | Behavior                                                                       |
| ------------------------------- | ---------------------------------- | ------------------------------------------------------------------------------ |
| `"none"`                        | `LLMModel.NO_TOOL_CHOICE`          | The model will not call any tools and instead generates a message              |
| `"auto"`                        | `LLMModel.MODEL_CHOOSES_TOOL`      | The model can choose between generating a message or calling one or more tools |
| `"required"`                    | `LLMModel.TOOL_CHOICE_REQUIRED`    | The model must call one or more tools                                          |
| A specific `aviary.Tool` object | N/A                                | The model must call this specific tool                                         |
| `None`                          | `LLMModel.UNSPECIFIED_TOOL_CHOICE` | No tool choice preference is provided to the LLM API                           |

When tools are provided, the LLM's response will be wrapped in a `ToolRequestMessage` instead of a regular `Message`. The key differences are:

* `Message` represents a basic chat message with a role (system/user/assistant) and content
* `ToolRequestMessage` extends `Message` to include `tool_calls`, which contains a list of `ToolCall` objects, which contains the tools the LLM chose to invoke and their arguments

Further details about how to define a tool, use the `ToolRequestMessage` and the `ToolCall` objects can be found in the [Aviary documentation](https://github.com/Future-House/aviary?tab=readme-ov-file#tool).

Here is a minimal example usage:

```python
from lmi import LiteLLMModel
from aviary.core import Message, Tool
import operator


# Define a function that will be used as a tool
def calculator(operation: str, x: float, y: float) -> float:
    """
    Performs basic arithmetic operations on two numbers.

    Args:
        operation (str): The arithmetic operation to perform ("+", "-", "*", or "/")
        x (float): The first number
        y (float): The second number

    Returns:
        float: The result of applying the operation to x and y

    Raises:
        KeyError: If operation is not one of "+", "-", "*", "/"
        ZeroDivisionError: If operation is "/" and y is 0
    """
    operations = {
        "+": operator.add,
        "-": operator.sub,
        "*": operator.mul,
        "/": operator.truediv,
    }
    return operations[operation](x, y)


# Create a tool from the calculator function
calculator_tool = Tool.from_function(calculator)

# The LLM must use the calculator tool
llm = LiteLLMModel()
result = await llm.call_single(
    messages=[Message(content="What is 2 + 2?")],
    tools=[calculator_tool],
    tool_choice=LiteLLMModel.TOOL_CHOICE_REQUIRED,
)

# result.messages[0] will be a ToolRequestMessage with tool_calls containing
# the calculator invocation with x=2, y=2, operation="+"
```

### Vertex

Vertex requires a bit of extra set-up. First, install the extra dependency for auth:

```sh
pip install google-api-python-client
```

and then you need to configure which region/project you're using for the model calls. Make sure you're authed for that region/project. Typically that means running:

```sh
gcloud auth application-default login
```

Then you can use vertex models:

```py
from lmi import LiteLLMModel
from aviary.core import Message

vertex_config = {"vertex_project": "PROJECT_ID", "vertex_location": "REGION"}

llm = LiteLLMModel(name="vertex_ai/gemini-2.5-pro", config=vertex_config)
await llm.call_single("hey")
```

### Embedding models

This client also includes embedding models. An embedding model is a class that inherits from `EmbeddingModel` and implements the `embed_documents` method, which receives a list of strings and returns a list with a list of floats (the embeddings) for each string.

Currently, the following embedding models are supported:

* `LiteLLMEmbeddingModel`
* `SparseEmbeddingModel`
* `SentenceTransformerEmbeddingModel`
* `HybridEmbeddingModel`

#### LiteLLMEmbeddingModel

`LiteLLMEmbeddingModel` provides a wrapper around LiteLLM's embedding functionality. It supports various embedding models through the LiteLLM interface, with automatic dimension inference and token limit handling. It defaults to `text-embedding-3-small` and can be configured with `name` and `config` parameters. Notice that `LiteLLMEmbeddingModel` can also be rate limited.

```python
from lmi import LiteLLMEmbeddingModel

model = LiteLLMEmbeddingModel(
    name="text-embedding-3-small",
    config={"rate_limit": "100/minute", "batch_size": 16},
)

embeddings = await model.embed_documents(["text1", "text2", "text3"])
```

#### HybridEmbeddingModel

`HybridEmbeddingModel` combines multiple embedding models by concatenating their outputs. It is typically used to combine a dense embedding model (like `LiteLLMEmbeddingModel`) with a sparse embedding model for improved performance. The model can be created in two ways:

```python
from lmi import LiteLLMEmbeddingModel, SparseEmbeddingModel, HybridEmbeddingModel

dense_model = LiteLLMEmbeddingModel(name="text-embedding-3-small")
sparse_model = SparseEmbeddingModel()
hybrid_model = HybridEmbeddingModel(models=[dense_model, sparse_model])
```

The resulting embedding dimension will be the sum of the dimensions of all component models. For example, if you combine a 1536-dimensional dense embedding with a 256-dimensional sparse embedding, the final embedding will be 1792-dimensional.

#### SentenceTransformerEmbeddingModel

You can also use `sentence-transformer`, which is a local embedding library with support for HuggingFace models, by installing `lmi[local]`.


# src


# ldp


# nn


# Chat Templates

* llama3.1\_chat\_template\_vllm.jinja: <https://github.com/vllm-project/vllm/blob/4fb8e329fd6f51d576bcf4b7e8907e0d83c4b5cf/examples/tool_chat_template_llama3.1_json.jinja>
* llama3.1\_chat\_template\_hf.jinja: <https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct/blob/main/tokenizer_config.json#L2053>
* llama3.1\_chat\_template\_thought.jinja: Fixes a typo in llama3.1\_chat\_template\_ori.jinja
* llama3.1\_chat\_template\_nothought.jinja: Derived from llama3.1\_chat\_template\_ori.jinja, but removes thoughts
* llama\*ori.jinja: TODOC - these were written by @kwanUm


