Python API
Data Types
Trajectory
This class records trajectory information from agent runs.
@dataclass
class Trajectory:
prompt_tokens: torch.Tensor
response_tokens: torch.Tensor
response_masks: torch.Tensor
idx: int = 0
trajectory_reward: float | int = 0.0
chat_completions: list[dict[str, str]] = field(default_factory=list)
metrics: dict[str, Any] = field(default_factory=lambda: {"steps": 0,
"toolcall_reward": 0.0,
"res_reward": 0.0,
"reward_time": 0.0,
"env_time": 0.0,
"llm_time": 0.0,
"total_time": 0.0})
| Parameter | Type | Description |
|---|---|---|
| idx | int | Trajectory index or ID. The default value is 0. |
| prompt_tokens | torch.Tensor | Token sequence of the input prompt. It cannot contain NaN or Inf. |
| response_tokens | torch.Tensor | Token sequence of the model response. It cannot contain NaN or Inf. |
| response_masks | torch.Tensor | Mask for response tokens, used to mark valid tokens. It cannot contain NaN or Inf. |
| trajectory_reward | float | int | Reward value of the trajectory. The default value is 0.0. |
| chat_completions | list[dict[str,str]] | List of LLM conversations. The default value is an empty list. |
| metrics | dict[str, Any] | Performance metrics of the trajectory. The Any values are as follows: steps: Total number of execution steps in the trajectory. The value is an integer. The default value is 0. reward_time: Time spent calculating the reward value. The value is of the numeric type. The default value is 0.0. toolcall_reward: Reward value from tool calls during trajectory generation. The value is of the numeric type. The default value is 0.0. res_reward: Reward value of the final answer. The value is of the numeric type. The default value is 0.0. env_time: Time spent interacting with the environment. The value is of the numeric type. The default value is 0.0. llm_time: Time spent on LLM inference. The value is of the numeric type. The default value is 0.0. total_time: Total time spent executing the entire trajectory. The value is of the numeric type. The default value is 0.0. |
StepTrajectory
This class records trajectory information for agent runs in step-level mode.
@dataclass
class Step:
chat_completions: list[dict[str, str]] = field(default_factory=list)
thought: str = ""
action: Any = None
observation: Any = None
model_response: str = ""
info: dict = field(default_factory=dict)
reward: float = 0.0
done: bool = False
mc_return: float = 0.0
@dataclass
class StepTrajectory(Trajectory):
task: Any = None
steps: list[Step] = field(default_factory=list)
Step class parameters
| Parameter | Type | Description |
|---|---|---|
| chat_completions | list[dict[str, str]] | Complete conversation context for the entire inference process, including historical turns. Used to build the model input. |
| thought | str | Content inside the <think> tag in the model response. This represents the model internal reasoning in the current step. |
| action | Any | Content inside the <tool call> tag in the model response. This represents the action the model chooses to take, such as a tool call. |
| observation | Any | External observation received in this step. For turn 0, this is the user's original question. For later turns, it is the result of the previous action, such as a tool return. |
| model_response | str | Complete response generated by the LLM, that is, the content of 'role': 'assistant'. |
| info | dict | Additional information dictionary. The default value is empty, and you can use it to record metadata such as tool IDs and time spent. |
| reward | float | Immediate reward obtained in this step. The default value is 0.0, which reflects the quality of the current action. |
| done | bool | Indicates whether the trajectory ends in this step. The default value is False, which marks whether the task is complete. |
| mc_return | float | Monte Carlo return from this step onward. The default value is 0.0, which is used for policy gradient training. |
StepTrajectory class parameters
| Parameter | Type | Description |
|---|---|---|
| task | Any | Original task input, such as the user query. The default value is None, and it serves as the initial goal of the entire trajectory. |
| steps | list[Step] | The ith Step contains the complete conversation context from turn 1 to turn i+1. |
Functional Functions
MemoryConfig
Class Description
The MemoryConfig class manages memory configuration, including thought content elision, summary generation, context window management, and model endpoints for chat and embeddings.
| Parameter | Type | Description | Value |
|---|---|---|---|
| simplify_thinking | bool | Indicates whether to simplify thinking content in messages. | The default value is False. |
| use_summary | bool | Enables automatic summarization of conversation history. | The default value is False. |
| max_summary_length | int | Maximum length of the generated summary, in tokens. | The default value is 1024. The value range is [1, max_prompt_length]. |
| max_prompt_length | int | Maximum total length of the prompt, including context, in tokens. | The default value is 8192. The value range is [1, 128K]. |
| before_raw_message | int | Number of initial messages at the beginning to keep unchanged. Messages in this range are not affected by summary generation or thought content elision. | The value must be greater than or equal to 0. The default value is 0. |
| end_raw_message | int | Number of initial messages at the end to keep unchanged. Messages in this range are not affected by summary generation or thought content elision. | The value must be less than or equal to 0. The default value is 0. |
| summary_system_prompt | str | System prompt template used to generate summaries. | The value is a fixed string. The specific value is shown in the following example. |
| oai_client | OpenAI | OpenAI client used to generate summaries. | The default value is empty. |
| oai_model_name | str | OpenAI model name used to generate summaries. | The default value is qwen2.5-7b-instruct. |
| train_model_tokenizer_path | str | Tokenizer path used to calculate context size. | The default value is an empty string. |
| model_config | dict | Applies Pydantic validation during assignment. validate_assignment=True ensures object integrity after creation by checking whether assigned values comply with field types and constraints. arbitrary_types_allowed=True means that the model allows arbitrary types as field values. |
The default value is {"validate_assignment": True, "arbitrary_types_allowed": True}. |
The summary_system_prompt template is as follows:
Please act as a summarization assistant to provide a precise and comprehensive summary of the specified content. The summary should meet the following requirements:
1. **Core Objective**: Extract the core information of the content, retain all key details (such as important data, viewpoints, time, characters, conclusions, etc.), and remove redundant information to ensure the summary is both concise and complete.
2. **Structural Requirements**:
- Begin with 1-2 sentences summarizing the content theme or core conclusion;
- List key information (such as main viewpoints, important events, data indicators, etc.) in bullet points, each containing specific details (avoid general statements);
- Conclude with the impact of the content, follow-up recommendations, or unresolved issues (if applicable).
3. **Detail Specifications**:
- Retain key terms, numerical values, and proper nouns (such as names of people, places, institutions) from the original text without alteration or abbreviation;
- If the content involves a timeline, organize key events in chronological order;
- If the content includes multiple viewpoints, clearly distinguish the positions of different entities;
- Avoid adding personal interpretations to maintain the objectivity of the summary.
4. **Length Recommendation**: The summary should not exceed {max_summary_length} tokens.
Please summarize the following dialogue content based on the above requirements.
MemorySummary
Class Description
The MemorySummary class provides memory management with automatic conversation summarization. It inherits from MemorySimple and automatically summarizes the conversation when the context length exceeds the configured limit.
| Parameter | Type | Description | Value |
|---|---|---|---|
| chat_client | SummaryClient | OpenAI client used to generate summaries. | The default value is empty. |
get_prompt_messages
Returns formatted messages for prompt generation and automatic summarization.
get_prompt_messages(config: dict | None = None) -> List[dict]
| Parameter | Type | Mandatory/Optional | Description |
|---|---|---|---|
| config | dict | Optional | Configuration parameter |
The get_prompt_messages method defined in the MemorySummary class takes config, which is the configuration dictionary used to update object attributes.
List[dict] is a non-empty list. Each element in the list is a dictionary in the form of {"role": "user", "content": "hello"}. The role value can be one of ["system", "user", "assistant", "tool", "summary"], and content must be a non-empty string.
The following is an example.
| Key | Description |
|---|---|
| role | Role that produces the context |
| content | Specific context information |
class MemorySummary(MemorySimple):
@validate_params(
config=dict(
validator=lambda x: isinstance(x, dict) or x is None, message="config must be a dictionary or None"
)
)
def get_prompt_messages(self, config: dict | None = None) -> list[dict]:
if config is not None:
self.update_config(config)
messages = self._get_effective_messages()
if self.config.use_summary and self._is_overlength(messages):
logger.info("Prompt length exceeds max_prompt_length, triggering summarization.")
messages = self._handle_overlength()
final_messages = self._format_summary_message(messages)
final_messages = self._apply_thinking_filter(final_messages)
if self._is_overlength(final_messages):
logger.warning(
f"PROMPT_TRUNCATION: Prompt length ({self._get_total_length(final_messages)}) "
f"exceeds max_prompt_length ({self.config.max_prompt_length})"
)
return MemorySimple._remove_message_other_key(final_messages)
BaseEngineWrapper
Class Description
The BaseEngineWrapper class provides a unified abstract interface that allows different AgentEngine implementations to adapt to it, allowing the AgenticRL framework to integrate with multiple AgentEngine types.
| Parameter | Type | Description | Value |
|---|---|---|---|
| agent_name | str | Name of the Agent scenario | The string can contain only letters, digits, and underscores (_), cannot start with a digit, and cannot be empty. |
| tokenizer | object | Text tokenizer object | The value cannot be empty and must be of type PreTrainedTokenizer or PreTrainedTokenizerFast. |
| sampling_params | dict | Sampling parameters used during model inference | The default value is empty. |
| max_prompt_length | int | Maximum length of the input prompt | The default value is 128K. The value range is [1, 128K]. |
| max_response_length | int | Maximum length of the output response | The default value is 8K. The value range is [1, 8K]. |
| n_parallel_agents | int | Number of agents executed in parallel | The default value is 8. The value range is [1, 64]. |
| max_steps | int | Maximum number of steps the agent can take | The default value is 5. The value range is [1, 10]. |
server_addresses
Description
The BaseEngineWrapper class provides local IP address strings for inference services. This property is a list, and each element is the IP address string of a vLLM inference service.
Property Description
| Property | Type | Mandatory/Optional | Description | Value Range |
|---|---|---|---|---|
| server_addresses | List[str] | Mandatory | Each element is the IP address of a vLLM inference service, in IP_ADDRESS:PORT format.These IP addresses can be used to access OpenAI-compatible RESTful APIs, including v1/completions and v1/chat/completions. The request and response formats of the APIs use JSON compatible with the OpenAI format. |
IP_ADDRESS:127.0.0.1PORT:0-65535. |
v1/completions
Parameters
| Parameter | Type | Mandatory/Optional | Description | Value Range |
|---|---|---|---|---|
| prompt | str | Mandatory | Input prompt for the LLM. | The value is a non-empty string. |
| n | int | Optional | Number of rollouts. | The value must be greater than or equal to 1. The default value is 1. |
| top_k | int | Optional | Sampling strategy parameter during text generation. It means that at each prediction step, the model considers only the top k candidate tokens with the highest probability. |
The value must be greater than or equal to 0. The default value is 50. |
| logprobs | int | Optional | Log probabilities of the n candidate tokens with the highest probabilities. |
The value must be greater than or equal to 0. The default value is 1. |
| min_p | float | Optional | Minimum probability threshold used to control the lower bound of candidate token probabilities during text generation. This threshold can be used to: min_p. This ensures that the generated text keeps only candidate tokens whose probability is above the threshold. min_p to control the determinism or diversity of the generated result. |
The default value is 0.0. The value range is [0.0, 1.0]. |
| detokenize | bool | Optional | Indicates whether to perform detokenization. | The default value is False, indicating that detokenization is not performed. The value can be False or True. |
| frequency_penalty | float | Optional | Frequency penalty. A penalty proportional to the number of occurrences is applied for frequent candidate tokens across the entire generated text. | The default value is 0.0. The value range is [-2.0, 2.0]. |
| max_tokens | int | Optional | Maximum number of candidate tokens allowed to be generated. The actual value is subject to configured limits. | The default value is 128. The value range is [1, 64000]. |
| min_tokens | int | Optional | Minimum number of candidate tokens required for each output sequence. The actual value is subject to configured limits. | The default value is 0. The value range is [0, 64000]. |
| presence_penalty | float | Optional | Presence penalty. A fixed penalty is applied to all tokens that have already appeared, regardless of frequency. | The default value is 0.0. The value range is [-2.0, 2.0]. |
| seed | int | Optional | Random seed for generating deterministic output. | The value range is [-65535, 65535]. |
| temperature | float | Optional | Sampling temperature. Higher values result in more random output. | The default value is 0.2. The value range is [0.0, 2.0]. |
| top_p | float | Optional | Nucleus sampling. Only candidate tokens with a cumulative probability mass within top_p are considered. |
[1e-8, 1.0] |
Example
import aiohttp
base_url = f"http://{self.server_addresses[0]}/v1/completions" # Use the v1/completions API.
headers = {"Content-Type": "application/json"}
async with aiohttp.ClientSession() as session:
async with session.post(base_url, headers=headers, json=completions_request) as response:
if response.status != 200:
err_msg = await response.text()
raise Exception(f"Http request failed. status = {response.status}: {err_msg}")
result = await response.json()
v1/chat/completions
Parameters
| Parameter | Type | Mandatory/Optional | Description | Value Range |
|---|---|---|---|---|
| messages | list[dict[str,str]] | Mandatory | Input prompt for the LLM. | The value is a non-empty list. Each element in the list is a dictionary in the form of {"role": "user", "content": "hello"}. The role value can be one of ["system", "user", "assistant", "tool"], and content must be a non-empty string. |
| n | int | Optional | Number of rollouts. | The value must be greater than or equal to 1. The default value is 1. |
| top_k | int | Optional | Sampling strategy parameter during text generation. It means that at each prediction step, the model considers only the top k candidate tokens with the highest probability. |
The value must be greater than or equal to 0. The default value is 50. |
| logprobs | int | Optional | Log probabilities of the n candidate tokens with the highest probabilities. |
The value must be greater than or equal to 0. The default value is 1. |
| min_p | float | Optional | Minimum probability threshold used to control the lower bound of candidate token probabilities during text generation. This threshold can be used to: min_p. This ensures that the generated text keeps only candidate tokens whose probability is above the threshold. min_p to control the determinism or diversity of the generated result. |
The default value is 0.0. The value range is [0.0, 1.0]. |
| detokenize | bool | Optional | Indicates whether to perform detokenization. | The default value is False, indicating that detokenization is not performed. The value can be False or True. |
| frequency_penalty | float | Optional | Frequency penalty. A penalty proportional to the number of occurrences is applied for frequent candidate tokens across the entire generated text. | The default value is 0.0. The value range is [-2.0, 2.0]. |
| max_tokens | int | Optional | Maximum number of candidate tokens allowed to be generated. The actual value is subject to configured limits. | The default value is 128. The value range is [1, 64000]. |
| min_tokens | int | Optional | Minimum number of candidate tokens required for each output sequence. The actual value is subject to configured limits. | The default value is 0. The value range is [0, 64000]. |
| presence_penalty | float | Optional | Presence penalty. A fixed penalty is applied to all tokens that have already appeared, regardless of frequency. | The default value is 0.0. The value range is [-2.0, 2.0]. |
| seed | int | Optional | Random seed for generating deterministic output. | The value range is [-65535, 65535]. |
| temperature | float | Optional | Sampling temperature. Higher values result in more random output. | The default value is 0.2. The value range is [0.0, 2.0]. |
| top_p | float | Optional | Nucleus sampling. Only candidate tokens with a cumulative probability mass within top_p are considered. |
[1e-8, 1.0] |
Example
import aiohttp
base_url = f"http://{self.server_addresses[0]}/v1/chat/completions" # Use the v1/chat/completions API.
headers = {"Content-Type": "application/json"}
async with aiohttp.ClientSession() as session:
async with session.post(base_url, headers=headers, json=completions_request) as response:
if response.status != 200:
err_msg = await response.text()
raise Exception(f"Http request failed. status = {response.status}: {err_msg}")
result = await response.json()
completions
The BaseEngineWrapper class provides inference capability for users. This property is a list of vLLM inference functions.
result = self.completions[0](completions_request)
| Parameter | Type | Mandatory/Optional | Description | Value Range |
|---|---|---|---|---|
| prompt | str | Mandatory | Input prompt for the LLM. | The value is a non-empty string. |
| n | int | Optional | Number of rollouts. | The value must be greater than or equal to 1. The default value is 1. |
| top_k | int | Optional | Sampling strategy parameter during text generation. It means that at each prediction step, the model considers only the top k candidate tokens with the highest probability. |
The value must be greater than or equal to 0. The default value is 50. |
| logprobs | int | Optional | Log probabilities of the n candidate tokens with the highest probabilities. |
The value must be greater than or equal to 0. The default value is 1. |
| min_p | float | Optional | Minimum probability threshold used to control the lower bound of candidate token probabilities during text generation. This threshold can be used to: min_p. This ensures that the generated text keeps only candidate tokens whose probability is above the threshold. min_p to control the determinism or diversity of the generated result. |
The default value is 0.0. The value range is [0.0, 1.0]. |
| detokenize | bool | Optional | Indicates whether to perform detokenization. | The default value is False, indicating that detokenization is not performed. The value can be False or True. |
| frequency_penalty | float | Optional | Frequency penalty. A penalty proportional to the number of occurrences is applied for frequent candidate tokens across the entire generated text. | The default value is 0.0. The value range is [-2.0, 2.0]. |
| max_tokens | int | Optional | Maximum number of candidate tokens allowed to be generated. The actual value is subject to configured limits. | The default value is 128. The value range is [1, 64000]. |
| min_tokens | int | Optional | Minimum number of candidate tokens required for each output sequence. The actual value is subject to configured limits. | The default value is 0. The value range is [0, 64000]. |
| presence_penalty | float | Optional | Presence penalty. A fixed penalty is applied to all tokens that have already appeared, regardless of frequency. | The default value is 0.0. The value range is [-2.0, 2.0]. |
| seed | int | Optional | Random seed for generating deterministic output. | The value range is [-65535, 65535]. |
| temperature | float | Optional | Sampling temperature. Higher values result in more random output. | The default value is 0.2. The value range is [0.0, 2.0]. |
| top_p | float | Optional | Nucleus sampling. Only candidate tokens with a cumulative probability mass within top_p are considered. |
[1e-8, 1.0] |
initialize
Performs the initialization process required by AgentEngine. The derived class implements the specific behavior.
initialize()
generate_agent_trajectories_async
Generates agent trajectories asynchronously with the agent execution engine.
generate_agent_trajectories_async(tasks: List[dict]) -> List[Trajectory]
| Parameter | Type | Mandatory/Optional | Description |
|---|---|---|---|
| tasks | List[dict] | Mandatory | Task list |
The abstract method generate_agent_trajectories_async defined by the abstract class BaseEngineWrapper provides a unified interface template. The specific implementation is defined by each derived class according to its own characteristics. The input parameter tasks is the task list constructed by Agent SDK. Each dictionary element in the list contains the following fields.
| Key | Description |
|---|---|
| id | Task ID |
| question | Question content of the current task |
| ground_truth | Correct answer for the task |
| Type | Description |
|---|---|
| List[Trajectory] | Sequence of generated agent trajectories. Each element is an object of type Trajectory. |
from agentic_rl import BaseEngineWrapper, Trajectory
class MockEngineWrapper(BaseEngineWrapper):
def initialize(self):
print("Initializing mock engine...")
def generate_agent_trajectories_async(self, tasks):
return [Trajectory(idx=0, prompt_tokens=torch.tensor([7, 8, 9]), response_tokens=torch.tensor([10, 11, 12]),
response_masks=torch.tensor([1, 1, 1]), trajectory_reward=2.0,
chat_completions=[{"role": "assistant", "content": "test"}],
metrics={"steps": 1, "reward_time": 2.0, "env_time": 3.0, "llm_time": 4.0, "total_time": 9.0})]
MockEngine = MockEngineWrapper(agent_name="mock_agent_name",
tokenizer=..., # Text tokenizer object
sampling_params={"mock":"sampling_params"},
max_prompt_length=128*1024,
max_response_length=8*1024,
n_parallel_agents=16,
max_steps=8)
trajectories = MockEngine.generate_agent_trajectories_async()