Python API

Data Types

Trajectory

Description

This class records trajectory information from agent runs.

Function prototype

@dataclass
class Trajectory:
    prompt_tokens: torch.Tensor
    response_tokens: torch.Tensor
    response_masks: torch.Tensor
    idx: int = 0
    trajectory_reward: float | int = 0.0
    chat_completions: list[dict[str, str]] = field(default_factory=list)
    metrics: dict[str, Any] = field(default_factory=lambda: {"steps": 0,
                                                             "toolcall_reward": 0.0,
                                                             "res_reward": 0.0,
                                                             "reward_time": 0.0,
                                                             "env_time": 0.0,
                                                             "llm_time": 0.0,
                                                             "total_time": 0.0})

Parameters

Parameter Type Description
idx int Trajectory index or ID. The default value is 0.
prompt_tokens torch.Tensor Token sequence of the input prompt. It cannot contain NaN or Inf.
response_tokens torch.Tensor Token sequence of the model response. It cannot contain NaN or Inf.
response_masks torch.Tensor Mask for response tokens, used to mark valid tokens. It cannot contain NaN or Inf.
trajectory_reward float | int Reward value of the trajectory. The default value is 0.0.
chat_completions list[dict[str,str]] List of LLM conversations. The default value is an empty list.
metrics dict[str, Any] Performance metrics of the trajectory. The Any values are as follows:
  • steps: Total number of execution steps in the trajectory. The value is an integer. The default value is 0.
  • reward_time: Time spent calculating the reward value. The value is of the numeric type. The default value is 0.0.
  • toolcall_reward: Reward value from tool calls during trajectory generation. The value is of the numeric type. The default value is 0.0.
  • res_reward: Reward value of the final answer. The value is of the numeric type. The default value is 0.0.
  • env_time: Time spent interacting with the environment. The value is of the numeric type. The default value is 0.0.
  • llm_time: Time spent on LLM inference. The value is of the numeric type. The default value is 0.0.
  • total_time: Total time spent executing the entire trajectory. The value is of the numeric type. The default value is 0.0.
  • StepTrajectory

    Description

    This class records trajectory information for agent runs in step-level mode.

    Function prototype

    @dataclass
    class Step:
        chat_completions: list[dict[str, str]] = field(default_factory=list)
        thought: str = ""
        action: Any = None
        observation: Any = None
        model_response: str = ""
        info: dict = field(default_factory=dict)
        reward: float = 0.0
        done: bool = False
        mc_return: float = 0.0
    
    @dataclass
    class StepTrajectory(Trajectory):
        task: Any = None
        steps: list[Step] = field(default_factory=list)
    

    Parameters

    Step class parameters

    Parameter Type Description
    chat_completions list[dict[str, str]] Complete conversation context for the entire inference process, including historical turns. Used to build the model input.
    thought str Content inside the <think> tag in the model response. This represents the model internal reasoning in the current step.
    action Any Content inside the <tool call> tag in the model response. This represents the action the model chooses to take, such as a tool call.
    observation Any External observation received in this step. For turn 0, this is the user's original question. For later turns, it is the result of the previous action, such as a tool return.
    model_response str Complete response generated by the LLM, that is, the content of 'role': 'assistant'.
    info dict Additional information dictionary. The default value is empty, and you can use it to record metadata such as tool IDs and time spent.
    reward float Immediate reward obtained in this step. The default value is 0.0, which reflects the quality of the current action.
    done bool Indicates whether the trajectory ends in this step. The default value is False, which marks whether the task is complete.
    mc_return float Monte Carlo return from this step onward. The default value is 0.0, which is used for policy gradient training.

    StepTrajectory class parameters

    Parameter Type Description
    task Any Original task input, such as the user query. The default value is None, and it serves as the initial goal of the entire trajectory.
    steps list[Step] The ith Step contains the complete conversation context from turn 1 to turn i+1.

    Functional Functions

    MemoryConfig

    Class Description

    Description

    The MemoryConfig class manages memory configuration, including thought content elision, summary generation, context window management, and model endpoints for chat and embeddings.

    Parameters

    Parameter Type Description Value
    simplify_thinking bool Indicates whether to simplify thinking content in messages. The default value is False.
    use_summary bool Enables automatic summarization of conversation history. The default value is False.
    max_summary_length int Maximum length of the generated summary, in tokens. The default value is 1024. The value range is [1, max_prompt_length].
    max_prompt_length int Maximum total length of the prompt, including context, in tokens. The default value is 8192. The value range is [1, 128K].
    before_raw_message int Number of initial messages at the beginning to keep unchanged. Messages in this range are not affected by summary generation or thought content elision. The value must be greater than or equal to 0. The default value is 0.
    end_raw_message int Number of initial messages at the end to keep unchanged. Messages in this range are not affected by summary generation or thought content elision. The value must be less than or equal to 0. The default value is 0.
    summary_system_prompt str System prompt template used to generate summaries. The value is a fixed string. The specific value is shown in the following example.
    oai_client OpenAI OpenAI client used to generate summaries. The default value is empty.
    oai_model_name str OpenAI model name used to generate summaries. The default value is qwen2.5-7b-instruct.
    train_model_tokenizer_path str Tokenizer path used to calculate context size. The default value is an empty string.
    model_config dict Applies Pydantic validation during assignment. validate_assignment=True ensures object integrity after creation by checking whether assigned values comply with field types and constraints. arbitrary_types_allowed=True means that the model allows arbitrary types as field values. The default value is {"validate_assignment": True, "arbitrary_types_allowed": True}.

    The summary_system_prompt template is as follows:

    Please act as a summarization assistant to provide a precise and comprehensive summary of the specified content. The summary should meet the following requirements:
    1. **Core Objective**: Extract the core information of the content, retain all key details (such as important data, viewpoints, time, characters, conclusions, etc.), and remove redundant information to ensure the summary is both concise and complete.
    2. **Structural Requirements**:
    - Begin with 1-2 sentences summarizing the content theme or core conclusion;
    - List key information (such as main viewpoints, important events, data indicators, etc.) in bullet points, each containing specific details (avoid general statements);
    - Conclude with the impact of the content, follow-up recommendations, or unresolved issues (if applicable).
    3. **Detail Specifications**:
    - Retain key terms, numerical values, and proper nouns (such as names of people, places, institutions) from the original text without alteration or abbreviation;
    - If the content involves a timeline, organize key events in chronological order;
    - If the content includes multiple viewpoints, clearly distinguish the positions of different entities;
    - Avoid adding personal interpretations to maintain the objectivity of the summary.
    4. **Length Recommendation**: The summary should not exceed {max_summary_length} tokens.
    Please summarize the following dialogue content based on the above requirements.
    

    MemorySummary

    Class Description

    Description

    The MemorySummary class provides memory management with automatic conversation summarization. It inherits from MemorySimple and automatically summarizes the conversation when the context length exceeds the configured limit.

    Parameters

    Parameter Type Description Value
    chat_client SummaryClient OpenAI client used to generate summaries. The default value is empty.

    get_prompt_messages

    Description

    Returns formatted messages for prompt generation and automatic summarization.

    Function prototype

    get_prompt_messages(config: dict | None = None) -> List[dict]
    

    Parameters

    Parameter Type Mandatory/Optional Description
    config dict Optional Configuration parameter

    The get_prompt_messages method defined in the MemorySummary class takes config, which is the configuration dictionary used to update object attributes.

    Returns

    List[dict] is a non-empty list. Each element in the list is a dictionary in the form of {"role": "user", "content": "hello"}. The role value can be one of ["system", "user", "assistant", "tool", "summary"], and content must be a non-empty string. The following is an example.

    Key Description
    role Role that produces the context
    content Specific context information

    Examples

    class MemorySummary(MemorySimple):
    
        @validate_params(
            config=dict(
                validator=lambda x: isinstance(x, dict) or x is None, message="config must be a dictionary or None"
            )
        )
        def get_prompt_messages(self, config: dict | None = None) -> list[dict]:
    
            if config is not None:
                self.update_config(config)
    
            messages = self._get_effective_messages()
    
            if self.config.use_summary and self._is_overlength(messages):
                logger.info("Prompt length exceeds max_prompt_length, triggering summarization.")
                messages = self._handle_overlength()
    
            final_messages = self._format_summary_message(messages)
    
            final_messages = self._apply_thinking_filter(final_messages)
    
            if self._is_overlength(final_messages):
                logger.warning(
                    f"PROMPT_TRUNCATION: Prompt length ({self._get_total_length(final_messages)}) "
                    f"exceeds max_prompt_length ({self.config.max_prompt_length})"
                )
    
            return MemorySimple._remove_message_other_key(final_messages)
    

    BaseEngineWrapper

    Class Description

    Description

    The BaseEngineWrapper class provides a unified abstract interface that allows different AgentEngine implementations to adapt to it, allowing the AgenticRL framework to integrate with multiple AgentEngine types.

    Parameters

    Parameter Type Description Value
    agent_name str Name of the Agent scenario The string can contain only letters, digits, and underscores (_), cannot start with a digit, and cannot be empty.
    tokenizer object Text tokenizer object The value cannot be empty and must be of type PreTrainedTokenizer or PreTrainedTokenizerFast.
    sampling_params dict Sampling parameters used during model inference The default value is empty.
    max_prompt_length int Maximum length of the input prompt The default value is 128K. The value range is [1, 128K].
    max_response_length int Maximum length of the output response The default value is 8K. The value range is [1, 8K].
    n_parallel_agents int Number of agents executed in parallel The default value is 8. The value range is [1, 64].
    max_steps int Maximum number of steps the agent can take The default value is 5. The value range is [1, 10].

    server_addresses

    Description

    The BaseEngineWrapper class provides local IP address strings for inference services. This property is a list, and each element is the IP address string of a vLLM inference service.

    Property Description

    Property Type Mandatory/Optional Description Value Range
    server_addresses List[str] Mandatory Each element is the IP address of a vLLM inference service, in IP_ADDRESS:PORT format.
    These IP addresses can be used to access OpenAI-compatible RESTful APIs, including v1/completions and v1/chat/completions. The request and response formats of the APIs use JSON compatible with the OpenAI format.
    IP_ADDRESS:127.0.0.1
    PORT:0-65535.
    v1/completions

    Parameters

    Parameter Type Mandatory/Optional Description Value Range
    prompt str Mandatory Input prompt for the LLM. The value is a non-empty string.
    n int Optional Number of rollouts. The value must be greater than or equal to 1. The default value is 1.
    top_k int Optional Sampling strategy parameter during text generation. It means that at each prediction step, the model considers only the top k candidate tokens with the highest probability. The value must be greater than or equal to 0. The default value is 50.
    logprobs int Optional Log probabilities of the n candidate tokens with the highest probabilities. The value must be greater than or equal to 0. The default value is 1.
  • 0: Does not output log probabilities.
  • Other values: Outputs log probabilities.
  • min_p float Optional Minimum probability threshold used to control the lower bound of candidate token probabilities during text generation. This threshold can be used to:
  • Filter out candidate tokens whose probability is lower than min_p. This ensures that the generated text keeps only candidate tokens whose probability is above the threshold.
  • Adjust the value of min_p to control the determinism or diversity of the generated result.
  • The default value is 0.0. The value range is [0.0, 1.0].
    detokenize bool Optional Indicates whether to perform detokenization. The default value is False, indicating that detokenization is not performed. The value can be False or True.
    frequency_penalty float Optional Frequency penalty. A penalty proportional to the number of occurrences is applied for frequent candidate tokens across the entire generated text. The default value is 0.0. The value range is [-2.0, 2.0].
    max_tokens int Optional Maximum number of candidate tokens allowed to be generated. The actual value is subject to configured limits. The default value is 128. The value range is [1, 64000].
    min_tokens int Optional Minimum number of candidate tokens required for each output sequence. The actual value is subject to configured limits. The default value is 0. The value range is [0, 64000].
    presence_penalty float Optional Presence penalty. A fixed penalty is applied to all tokens that have already appeared, regardless of frequency. The default value is 0.0. The value range is [-2.0, 2.0].
    seed int Optional Random seed for generating deterministic output. The value range is [-65535, 65535].
    temperature float Optional Sampling temperature. Higher values result in more random output. The default value is 0.2. The value range is [0.0, 2.0].
    top_p float Optional Nucleus sampling. Only candidate tokens with a cumulative probability mass within top_p are considered. [1e-8, 1.0]

    Example

    import aiohttp
    
    base_url = f"http://{self.server_addresses[0]}/v1/completions" # Use the v1/completions API.
    headers = {"Content-Type": "application/json"}
    
    async with aiohttp.ClientSession() as session:
        async with session.post(base_url, headers=headers, json=completions_request) as response:
            if response.status != 200:
                err_msg = await response.text()
                raise Exception(f"Http request failed. status = {response.status}: {err_msg}")
            result = await response.json()
    
    v1/chat/completions

    Parameters

    Parameter Type Mandatory/Optional Description Value Range
    messages list[dict[str,str]] Mandatory Input prompt for the LLM. The value is a non-empty list. Each element in the list is a dictionary in the form of {"role": "user", "content": "hello"}. The role value can be one of ["system", "user", "assistant", "tool"], and content must be a non-empty string.
    n int Optional Number of rollouts. The value must be greater than or equal to 1. The default value is 1.
    top_k int Optional Sampling strategy parameter during text generation. It means that at each prediction step, the model considers only the top k candidate tokens with the highest probability. The value must be greater than or equal to 0. The default value is 50.
    logprobs int Optional Log probabilities of the n candidate tokens with the highest probabilities. The value must be greater than or equal to 0. The default value is 1.
  • 0: Does not output log probabilities.
  • Other values: Outputs log probabilities.
  • min_p float Optional Minimum probability threshold used to control the lower bound of candidate token probabilities during text generation. This threshold can be used to:
  • Filter out candidate tokens whose probability is lower than min_p. This ensures that the generated text keeps only candidate tokens whose probability is above the threshold.
  • Adjust the value of min_p to control the determinism or diversity of the generated result.
  • The default value is 0.0. The value range is [0.0, 1.0].
    detokenize bool Optional Indicates whether to perform detokenization. The default value is False, indicating that detokenization is not performed. The value can be False or True.
    frequency_penalty float Optional Frequency penalty. A penalty proportional to the number of occurrences is applied for frequent candidate tokens across the entire generated text. The default value is 0.0. The value range is [-2.0, 2.0].
    max_tokens int Optional Maximum number of candidate tokens allowed to be generated. The actual value is subject to configured limits. The default value is 128. The value range is [1, 64000].
    min_tokens int Optional Minimum number of candidate tokens required for each output sequence. The actual value is subject to configured limits. The default value is 0. The value range is [0, 64000].
    presence_penalty float Optional Presence penalty. A fixed penalty is applied to all tokens that have already appeared, regardless of frequency. The default value is 0.0. The value range is [-2.0, 2.0].
    seed int Optional Random seed for generating deterministic output. The value range is [-65535, 65535].
    temperature float Optional Sampling temperature. Higher values result in more random output. The default value is 0.2. The value range is [0.0, 2.0].
    top_p float Optional Nucleus sampling. Only candidate tokens with a cumulative probability mass within top_p are considered. [1e-8, 1.0]

    Example

    import aiohttp
    
    base_url = f"http://{self.server_addresses[0]}/v1/chat/completions" # Use the v1/chat/completions API.
    headers = {"Content-Type": "application/json"}
    
    async with aiohttp.ClientSession() as session:
        async with session.post(base_url, headers=headers, json=completions_request) as response:
            if response.status != 200:
                err_msg = await response.text()
                raise Exception(f"Http request failed. status = {response.status}: {err_msg}")
            result = await response.json()
    

    completions

    Description

    The BaseEngineWrapper class provides inference capability for users. This property is a list of vLLM inference functions.

    Function prototype

    result = self.completions[0](completions_request)
    

    Parameters

    Parameter Type Mandatory/Optional Description Value Range
    prompt str Mandatory Input prompt for the LLM. The value is a non-empty string.
    n int Optional Number of rollouts. The value must be greater than or equal to 1. The default value is 1.
    top_k int Optional Sampling strategy parameter during text generation. It means that at each prediction step, the model considers only the top k candidate tokens with the highest probability. The value must be greater than or equal to 0. The default value is 50.
    logprobs int Optional Log probabilities of the n candidate tokens with the highest probabilities. The value must be greater than or equal to 0. The default value is 1.
  • 0: Does not output log probabilities.
  • Other values: Outputs log probabilities.
  • min_p float Optional Minimum probability threshold used to control the lower bound of candidate token probabilities during text generation. This threshold can be used to:
  • Filter out candidate tokens whose probability is lower than min_p. This ensures that the generated text keeps only candidate tokens whose probability is above the threshold.
  • Adjust the value of min_p to control the determinism or diversity of the generated result.
  • The default value is 0.0. The value range is [0.0, 1.0].
    detokenize bool Optional Indicates whether to perform detokenization. The default value is False, indicating that detokenization is not performed. The value can be False or True.
    frequency_penalty float Optional Frequency penalty. A penalty proportional to the number of occurrences is applied for frequent candidate tokens across the entire generated text. The default value is 0.0. The value range is [-2.0, 2.0].
    max_tokens int Optional Maximum number of candidate tokens allowed to be generated. The actual value is subject to configured limits. The default value is 128. The value range is [1, 64000].
    min_tokens int Optional Minimum number of candidate tokens required for each output sequence. The actual value is subject to configured limits. The default value is 0. The value range is [0, 64000].
    presence_penalty float Optional Presence penalty. A fixed penalty is applied to all tokens that have already appeared, regardless of frequency. The default value is 0.0. The value range is [-2.0, 2.0].
    seed int Optional Random seed for generating deterministic output. The value range is [-65535, 65535].
    temperature float Optional Sampling temperature. Higher values result in more random output. The default value is 0.2. The value range is [0.0, 2.0].
    top_p float Optional Nucleus sampling. Only candidate tokens with a cumulative probability mass within top_p are considered. [1e-8, 1.0]

    initialize

    Description

    Performs the initialization process required by AgentEngine. The derived class implements the specific behavior.

    Function prototype

    initialize()
    

    generate_agent_trajectories_async

    Description

    Generates agent trajectories asynchronously with the agent execution engine.

    Function prototype

    generate_agent_trajectories_async(tasks: List[dict]) -> List[Trajectory]
    

    Parameters

    Parameter Type Mandatory/Optional Description
    tasks List[dict] Mandatory Task list

    The abstract method generate_agent_trajectories_async defined by the abstract class BaseEngineWrapper provides a unified interface template. The specific implementation is defined by each derived class according to its own characteristics. The input parameter tasks is the task list constructed by Agent SDK. Each dictionary element in the list contains the following fields.

    Key Description
    id Task ID
    question Question content of the current task
    ground_truth Correct answer for the task

    Returns

    Type Description
    List[Trajectory] Sequence of generated agent trajectories. Each element is an object of type Trajectory.

    Examples

    from agentic_rl import BaseEngineWrapper, Trajectory
    
    class MockEngineWrapper(BaseEngineWrapper):
        def initialize(self):
            print("Initializing mock engine...")
    
        def generate_agent_trajectories_async(self, tasks):
            return [Trajectory(idx=0, prompt_tokens=torch.tensor([7, 8, 9]), response_tokens=torch.tensor([10, 11, 12]),
                               response_masks=torch.tensor([1, 1, 1]), trajectory_reward=2.0,
                               chat_completions=[{"role": "assistant", "content": "test"}],
                               metrics={"steps": 1, "reward_time": 2.0, "env_time": 3.0, "llm_time": 4.0, "total_time": 9.0})]
    
    MockEngine = MockEngineWrapper(agent_name="mock_agent_name",
                                   tokenizer=...,               # Text tokenizer object
                                   sampling_params={"mock":"sampling_params"},
                                   max_prompt_length=128*1024,
                                   max_response_length=8*1024,
                                   n_parallel_agents=16,
                                   max_steps=8)
    trajectories = MockEngine.generate_agent_trajectories_async()