MiroThinker is a deep research agent optimized for complex research and prediction tasks. Our latest models, MiroThinker-1.7, achieves 74.0 and 75.3 on the BrowseComp and BrowseComp Zh, respectively.
MiroThinker:一款专为研究与预测任务优化的深度研究智能体。它在极具挑战性的 BrowseComp 基准测试中取得了 88.2 分的成绩。详见快速开始。
📋 目录
📰 新闻与更新
- [2026-03-11] 🎉🎉🎉 推出 MiroThinker-1.7,包括 MiroThinker-1.7-mini 和 MiroThinker-1.7。MiroThinker-1.7-mini 在 BrowseComp-ZH 上达到 72.3 分,以仅 300 亿参数的规模刷新了开源模型的性能纪录。我们的专有智能体 MiroThinker-H1 在 BrowseComp 和 BrowseComp-ZH 上的表现领先于所有开源及商业模型。
- [2026-01-23] 🎉 MiroThinker 在线平台 迎来两项重要更新:(a) 核心研究报告生成:深度研究在线报告现已支持生成、预览与分享功能。(b) 扩展文档上传类型:新增支持多种文件格式上传,如
.pdf、.doc、.ppt、.xls、.jpg。欢迎体验!MiroThinker 将持续维护并迭代升级,致力于成为您用过的最佳研究智能体! - [2026-01-05] 🎉🎉 我们发布了 MiroThinker-v1.5,这是一系列针对金融预测优化的开源深度研究智能体。MiroThinker-v1.5-30B 以更低成本超越了 Kimi-K2-Thinking 在 BrowseComp-ZH 上的表现,参数规模仅为其 1/30。MiroThinker-v1.5-235B 在 HLE-Text 上得分 39.2%,BrowseComp 上得分 69.8%,BrowseComp-ZH 上得分 71.5%,GAIA-Val-165 上得分 80.8%,刷新了搜索智能体的性能纪录。
📜 点击展开更早更新
- [2025-11-13] 🎉 MiroThinker-v1.0 正式发布!引入“交互式扩展”作为提升性能的第三维度,MiroThinker v1.0 支持 256K 上下文窗口及单任务最多 600 次工具调用。提供 80 亿、300 亿和 720 亿参数规模版本,在 HLE-Text、BrowseComp、BrowseComp-ZH 和 GAIA-Text-103 上分别取得 37.7%、47.1%、55.6% 和 81.9% 的成绩。更多详情参见技术报告。
- [2025-09-11] MiroThinker-72B-Preview 在本周 FutureX 基准测试中排名第四。详见 FutureX。
- [2025-09-08] MiroThinker-v0.2 现已发布,在多项基准测试中取得开源最佳性能,包括 HLE(17.8%)、HLE-Text-Only(19.1%)、BrowseComp-EN(17.2%)、BrowseComp-ZH(29.4%)、XBench-DeepSearch(56.0%)和 Frames(74.8%)。
- [2025-09-07] 我们支持了更多基准测试,包括 BrowseComp-ZH、XBench-DeepSearch 和 FutureX。未来计划添加更多基准测试。
- [2025-08-22] 推出 MiroThinker 简化部署方案,优化资源占用并加快启动时间。体验交互式演示:🚀 尝试 Gradio 演示
- [2025-08-08] MiroThinker-v0.1 发布。
📝 简介
MiroThinker-1.7
全新的MiroThinker系列在构建可靠的长链任务智能体方面实现了重大飞跃。通过增强的后训练流程精心打造,MiroThinker-1.7系列在开源模型中,于深度研究任务上取得了SOTA性能。
核心特性
- 🚀 MiroThinker-1.7支持256K上下文窗口、长周期推理及深度多步骤分析。
- 🔧 每个任务最多可处理300次工具交互,目前具备更精准的逐步推理与决策能力。
- 📦 提供30B和235B参数规模版本,并附带全面的工具套件与工作流,可灵活支持多样化的研究场景及计算预算。
- 我们的专有智能体MiroThinker-H1为长链可验证推理提供了令人鼓舞的证据——推理过程可逐步验证且全局可验证,提升了复杂智能体工作流的性能。
| 模型名称 | 参数规模 | 最大上下文 | 最大工具调用次数 | HF链接 |
|---|---|---|---|---|
| MiroThinker-1.7-mini | 30B | 256K | 300 | 🤗 link |
| MiroThinker-1.7 | 235B | 256K | 300 | 🤗 link |
MiroThinker-1.7在广泛的基准测试中展现出强大的通用研究性能,在BrowseComp、BrowseComp-ZH、GAIA-Val-165和HLE-Text上分别达到74.0%、75.3%、82.7%和42.9%。MiroThinker-1.7在BrowseComp-ZH上取得了SOTA性能。

MiroThinker-v1.5
📦 点击展开MiroThinker-v1.5详情
MiroThinker v1.5是全球领先的开源搜索智能体,它通过交互式扩展推动工具增强推理的发展——训练智能体处理更深层次、更频繁的智能体-环境交互,将其作为超越模型大小和上下文长度的第三个性能提升维度。

核心特性
- 🚀 MiroThinker v1.5支持256K上下文窗口、长周期推理及深度多步骤分析。
- 🔧 每个任务最多可处理400次工具调用——相比以往的开源研究智能体有显著提升。
- 📦 提供30B和235B参数规模版本,并附带全面的工具套件与工作流,可灵活支持多样化的研究场景及计算预算。
| 智能体名称 | 基础智能体 | 最大上下文 | 最大工具调用次数 | HF链接 |
|---|---|---|---|---|
| MiroThinker-v1.5-30B | Qwen3-30B-A3B-Thinking-2507 | 256K | 400 | 🤗 link |
| MiroThinker-v1.5-235B | Qwen3-235B-A22B-Thinking-2507 | 256K | 400 | 🤗 link |
MiroThinker v1.5在广泛的基准测试中展现出强大的通用研究性能,在HLE-Text、BrowseComp、BrowseComp-ZH和GAIA-Val-165上分别达到39.2%、69.8%、71.5%和80.8%。这些结果超越了以往的开源智能体,并创造了新的BrowseComp性能世界纪录。

MiroThinker-v1.0
📦 点击展开 MiroThinker-v1.0 详情
不同于以往仅通过模型规模或上下文长度进行扩展的智能体,MiroThinker v1.0 在智能体层面引入了交互式扩展,系统地训练智能体处理更深层次、更频繁的智能体-环境交互,将其作为提升性能的第三维度。交互式扩展利用环境反馈和外部信息获取来纠正错误并优化行动轨迹。

✨ 核心特性
- 🚀 256K 上下文窗口:支持长周期推理和深度多步骤分析
- 🔧 600 次工具调用:每个任务最多可处理 600 次工具调用,较以往开源研究智能体有显著提升
- 📦 多尺度版本:提供 8B、30B 和 72B 参数规模,配套全面的工具集和工作流,可灵活支持多样化的研究场景和计算预算
| 智能体名称 | 基础智能体 | 最大上下文 | 最大工具调用次数 | HF 链接 |
|---|---|---|---|---|
| MiroThinker-v1.0-8B | Qwen3-8B | 256K | 600 | 🤗 link |
| MiroThinker-v1.0-30B | Qwen3-30B-A3B-Thinking-2507 | 256K | 600 | 🤗 link |
| MiroThinker-v1.0-72B | Qwen2.5-72B-Instruct | 256K | 600 | 🤗 link |
MiroThinker v1.0 在各类基准测试中展现出强大的通用研究性能,在 HLE-Text、BrowseComp、BrowseComp-ZH 和 GAIA-Text-103 上的得分分别为37.7%、47.1%、55.6% 和 81.9%。这些结果超越了以往的开源智能体,并缩小了与 GPT-5-high 等商业智能体之间的差距。
MiroThinker-v0.2
📦 点击展开 MiroThinker-v0.2 详情
在这一新版本中,我们引入了三项关键改进:
- 📚 更丰富的训练数据,涵盖中英文来源,显著提升了基准测试性能和模型泛化能力
- 🎯 统一 DPO 训练,所有智能体共享单一偏好数据集
- 📏 扩展上下文长度,从 40k 增至 64k,以应对更具挑战性的多轮工具使用任务
与 v0.1 相比,MiroThinker v0.2 在各项基准测试中均实现了稳步提升。例如,在 GAIA-Text-103 上的得分从 57.3 提升至 64.1,在 BrowseComp-ZH 上从 17.0 提升至 29.4,这些数据充分体现了模型在通用研究智能体能力方面的显著进步。
| 智能体名称 | 基础智能体 | 最大上下文长度 | HF 链接 |
|---|---|---|---|
| MiroThinker-4B-SFT-v0.2 | Qwen3-4B | 64K | 🤗 link |
| MiroThinker-4B-DPO-v0.2 | Qwen3-4B | 64K | 🤗 link |
| MiroThinker-8B-SFT-v0.2 | Qwen3-8B | 64K | 🤗 link |
| MiroThinker-8B-DPO-v0.2 | Qwen3-8B | 64K | 🤗 link |
| MiroThinker-14B-SFT-v0.2 | Qwen3-14B | 64K | 🤗 link |
| MiroThinker-14B-DPO-v0.2 | Qwen3-14B | 64K | 🤗 link |
| MiroThinker-32B-SFT-v0.2 | Qwen3-32B | 64K | 🤗 link |
| MiroThinker-32B-DPO-v0.2 | Qwen3-32B | 64K | 🤗 link |
MiroThinker-v0.1
📦 点击展开 MiroThinker-v0.1 详情
开源智能体在 GAIA-Validation 基准测试中的性能表现。
我们已发布 MiroThinker v0.1 系列模型,包括参数规模为 8B、14B 和 32B 的 SFT 与 DPO 两种变体。值得关注的是,在用于评估高级智能体能力的严格评测套件 GAIA 基准 中,MiroThinker v0.1 取得了开源模型中的最先进性能,充分展现了其在长上下文、决策密集型及现实世界任务场景中的强大实力。
✨ 核心特性
🤖 MiroThinker 优化框架
- 🔓 全开源智能体框架:框架与智能体完全开源,确保透明度
- 🔗 工具集成:与外部工具及 API 无缝对接
- 📝 轨迹收集:全面记录并分析智能体交互过程,显示耗时及预计完成时间(以分钟为单位)。支持 SFT 和 DPO 训练
- 📊 基准评测:在多个基准数据集上进行广泛测试
📊 综合基准测试套件
📋 点击展开基准测试列表
- GAIA Validation:通用人工智能助手基准测试。(论文)
- GAIA-Text-103:GAIA Validation 的纯文本任务子集。(论文)
- HLE:人类终极考试。(论文)
- HLE-Text-2158:HLE 的纯文本任务子集。(论文)
- HLE-Text-500:HLE 的纯文本任务子集,由 WebThinker 创建。(论文)
- BrowseComp-EN:网络浏览与理解任务。(论文)
- BrowseComp-ZH:BrowseComp 的中文版本。(论文)
- WebWalkerQA:网络导航与问答。(论文)
- Frames:事实性、检索与推理测量集(Factuality, Retrieval, And reasoning MEasurement Set)。(论文)
- XBench-DeepSearch:深度研究智能体基准测试。(网站)
- FutureX:用于预测未知未来的实时基准测试。(网站)
- SEAL-0:评估大型语言模型在含冲突证据网络问题上表现的基准测试。(论文)
- AIME2025:2025 年美国数学邀请赛。(网站)
- DeepSearchQA:谷歌深度搜索问答基准测试。(论文)
📈 基准测试性能
MiroThinker-1.7
为防止潜在的信息泄露(例如从 HuggingFace 检索基准测试答案),我们在评估期间屏蔽了部分网站的访问权限。
MiroThinker-v1.5
📦 点击展开 MiroThinker-v1.5 详情
为防止潜在的信息泄露(例如从 HuggingFace 搜索基准测试答案),这些工具中已明确禁用对 HuggingFace 的访问。
我们进一步对所有轨迹的工具输出执行金丝雀字符串测试,并忽略任何被发现受到污染的轨迹,将其视为错误答案。
MiroThinker-v1.0
📦 点击展开 MiroThinker-v1.0 详情
MiroThinker-v0.2
MiroThinker-v0.1
📦 点击展开 MiroThinker-v0.1 详情
GAIA 基准测试
| 方法 | Text-103 最佳 Pass@1 |
Text-103 Pass@1(Avg@8) |
Val-165 最佳 Pass@1 |
Val-165 Pass@1(Avg@8) |
|---|---|---|---|---|
| 🔹—— 7B/8B 智能体 —— | ||||
| Search-o1-7B | 17.5 | - | - | - |
| R1-Searcher-7B | 20.4 | - | - | - |
| WebDancer-7B | 31.0 | - | - | - |
| WebSailor-7B | 37.9 | - | - | - |
| CK-Pro-8B | 40.3 | - | 32.7 | - |
| MiroThinker-8B-SFT-v0.1 | 44.7 | 40.1 | 34.6 | 31.8 |
| + 商业工具 | 46.6 | 42.1 | 37.6 | 33.9 |
| MiroThinker-8B-DPO-v0.1 | 46.6 | 44.8 | 37.0 | 35.4 |
| + 商业工具 | 50.5 | 46.7 | 38.2 | 35.9 |
| 🔹—— 14B 智能体 —— | ||||
| MiroThinker-14B-SFT-v0.1 | 47.6 | 44.4 | 37.0 | 34.4 |
| + 商业工具 | 49.5 | 47.5 | 41.8 | 39.8 |
| MiroThinker-14B-DPO-v0.1 | 48.5 | 46.6 | 42.4 | 39.2 |
| + 商业工具 | 52.4 | 48.5 | 45.5 | 42.0 |
| 🔹—— 32B 智能体 —— | ||||
| Qwen3-32B | 31.1 | 26.7 | 29.7 | 26.4 |
| Search-o1-32B | 28.2 | - | - | - |
| WebThinker-32B-RL | 48.5 | - | - | - |
| WebDancer-QwQ-32B | 51.5 | - | - | - |
| WebSailor-32B | 53.2 | - | - | - |
| WebShaper-QwQ-32B | 53.3 | - | - | - |
| MiroThinker-32B-SFT-v0.1 | 55.3 | 51.3 | 44.9 | 42.7 |
| + 商业工具 | 58.3 | 54.2 | 48.5 | 45.8 |
| MiroThinker-32B-DPO-v0.1 | 57.3 | 54.1 | 48.5 | 45.9 |
| + 商业工具 | 60.2 | 57.9 | 50.9 | 48.9 |
-
遵循 WebThinker、WebAgents 和 CognitiveKernel 的做法,我们报告三次运行中的最高得分 Best Pass@1,该得分通常反映更强的性能,尽管可能存在一定的波动性。为提供更稳定的衡量标准,我们额外报告 Pass@1(Avg@8),该指标以略低的分数为代价提供了更高的一致性。
-
为与先前的开源工作保持一致,我们使用 WebAgents LLM-as-a-Judge 模板评估 GAIA-Text-103,并使用官方 GAIA 评分脚本报告 GAIA-Val-165 的结果。
-
默认情况下,我们尽可能使用开源工具,但代码工具 E2B 和谷歌搜索工具 Serper 除外。我们在实现中使用了 Whisper、Qwen2.5-VL-72B-Instruct 和 Qwen3-235B-A22B-Thinking-2507。该框架可轻松扩展到您选择的其他开源工具。
-
将这些开源工具替换为商业替代方案可以带来性能提升。商业工具主要用于多模态能力和某些复杂推理子任务。大部分任务,包括规划、浏览、优化、导航等,均由我们的智能体处理。
更多基准测试
| 方法 | HLE Pass@1 |
Frames Pass@1 |
BrowseComp Pass@1 |
BrowseComp-ZH Pass@1 |
WebWalkerQA Pass@1 |
|---|---|---|---|---|---|
| OpenAI Deep Research | 26.6 | - | 51.5 | 42.9 | - |
| Gemini Deep Research | 26.9 | - | - | - | - |
| Kimi-Researcher | 26.9 | 78.8 | - | - | - |
| WebDancer-7B | - | - | - | - | 36.0 |
| WebSailor-7B | - | - | 6.7 | 14.2 | - |
| MiroThinker-8B-SFT-v0.1 | - | 58.0 | 5.5 | 9.3 | 41.3 |
| MiroThinker-8B-DPO-v0.1 | - | 64.4 | 8.7 | 13.6 | 45.7 |
| WebThinker-32B-RL | - | - | - | - | 46.5 |
| WebDancer-QwQ-32B | - | - | 3.8 | 18.0 | 47.9 |
| WebSailor-32B | - | - | 10.5 | 25.5 | - |
| WebShaper-32B | - | - | - | - | 51.4 |
| MiroThinker-32B-SFT-v0.1 | 10.2 | 70.4 | 10.6 | 13.8 | 45.7 |
| MiroThinker-32B-DPO-v0.1 | 11.8 | 71.7 | 13.0 | 17.0 | 49.3 |
-
MiroThinker 的性能通过本仓库和开源工具进行测试;其他智能体的结果来自其论文和官方网站。
-
由于 MiroVerse-v0.1 主要包含英文数据,智能体的中文能力有限。我们计划在下一版本中添加更多中文数据以提升性能。
🚀 快速开始
为获得最佳使用体验,我们建议在启用工具支持的智能体框架和思考模式的情况下使用 MiroThinker。
前提条件
- 🐍 Python 3.10 及以上版本
- 📦 uv 包管理器(安装指南)
- 🔑 所需 API 密钥(详见下文配置部分)
安装
# Clone the repository
git clone https://github.com/MiroMindAI/MiroThinker
cd MiroThinker
# Setup environment
cd apps/miroflow-agent
uv sync
# Configure API keys
cp .env.example .env
# Edit .env with your API keys (SERPER_API_KEY, JINA_API_KEY, E2B_API_KEY, etc.)
📝 环境变量:所需 API 密钥请参见工具配置部分。
工具配置
MiroThinker-1.7 的最小配置
| 服务器 | 描述 | 提供的工具 | 所需环境变量 |
|---|---|---|---|
tool-python |
执行环境和文件管理(E2B 沙箱) | create_sandbox、run_command、run_python_code、upload_file_from_local_to_sandbox、download_file_from_sandbox_to_local、download_file_from_internet_to_sandbox |
E2B_API_KEY |
search_and_scrape_webpage |
通过 Serper API 进行谷歌搜索 | google_search |
SERPER_API_KEY、SERPER_BASE_URL |
jina_scrape_llm_summary |
基于 LLM 的网页抓取与信息提取 | scrape_and_extract_info |
JINA_API_KEY、JINA_BASE_URL、SUMMARY_LLM_BASE_URL、SUMMARY_LLM_MODEL_NAME、SUMMARY_LLM_API_KEY |
最小 .env 配置示例:
# Required for MiroThinker v1.5 and v1.0 (minimal setup)
SERPER_API_KEY=your_serper_key
SERPER_BASE_URL="https://google.serper.dev"
JINA_API_KEY=your_jina_key
JINA_BASE_URL="https://r.jina.ai"
E2B_API_KEY=your_e2b_key
# Required for jina_scrape_llm_summary
# Note: Summary LLM can be a small model (e.g., Qwen3-14B or GPT-5-Nano)
# The choice has minimal impact on performance, use what's most convenient
SUMMARY_LLM_BASE_URL="https://your_summary_llm_base_url/v1/chat/completions"
SUMMARY_LLM_MODEL_NAME=your_llm_model_name # e.g., "Qwen/Qwen3-14B" or "gpt-5-nano"
SUMMARY_LLM_API_KEY=your_llm_api_key # Optional, depends on LLM provider
# Required for benchmark evaluation (LLM-as-a-Judge)
OPENAI_API_KEY=your_openai_key # Required for running benchmark evaluations
OPENAI_BASE_URL="https://api.openai.com/v1" # Optional, defaults to OpenAI's API
💡 为何如此精简:这 3 个 MCP 服务器涵盖了研究任务所需的核心能力:网络搜索、内容提取和代码执行。其他所有服务器均为可选增强功能。
🤖 摘要 LLM:
SUMMARY_LLM可以是像 Qwen3-14B 或 GPT-5-Nano 这样的小型模型。该选择对整体性能影响极小,使用最适合您设置的模型即可。📊 用于基准测试评估:如果您计划运行基准测试评估,还需要
OPENAI_API_KEY(以及可选的OPENAI_BASE_URL),以便在评估脚本中使用 LLM 作为评判器(LLM-as-a-Judge)功能。🖼️ 用于 GAIA 多模态任务:GAIA-Val-165 包含涉及图像/音频/视频文件的任务。由于 MiroThinker 是纯文本 LLM,因此使用 GPT-4o 将这些文件预处理为文本描述。相同的
OPENAI_API_KEY既用于此预处理,也用于 LLM 作为评判器。📖 更多详情:有关所有可用工具的完整文档,请参见 MiroFlow 工具 README。
🔧 点击展开其他可用工具
以下为可选工具,但未在 MiroThinker v1.0-1.7 评估中使用:
| 服务器名称 | 类型 | 描述 |
|---|---|---|
tool-vqa |
商业 | 使用 Claude 进行视觉处理 |
tool-vqa-os |
开源 | 视觉处理(开源替代方案) |
tool-transcribe |
商业 | 使用 OpenAI 进行音频转录 |
tool-transcribe-os |
开源 | 使用 Whisper 进行音频转录 |
tool-reasoning |
商业 | 使用 Claude 的推理引擎 |
tool-reasoning-os |
开源 | 推理引擎(开源替代方案) |
tool-reading |
开源 | 使用 MarkItDown 进行文档阅读 |
tool-google-search |
商业 | 使用 Google + 网页抓取进行网络搜索 |
tool-sogou-search |
商业 | 使用搜狗进行网络搜索(中文) |
📖 本地部署:有关在本地部署开源工具(
tool-vqa-os、tool-transcribe-os、tool-reasoning-os)的说明,请参见 本地工具部署指南。
有关所有可用工具的完整文档,请参见 MiroFlow 工具 README。
预配置代理设置
apps/miroflow-agent/conf/agent/ 目录包含多个预配置的代理设置。每个配置使用不同的工具,并要求在您的 .env 文件中设置相应的环境变量。
💡 推荐:对于 MiroThinker-1.7,使用
mirothinker_1.7_keep5_max200(带有上下文管理,推荐用于大多数任务)或mirothinker_v1.7_keep5_max300(仅用于 BrowseComp 和 BrowseComp-ZH)。
| 配置名称 | 描述 | 最大轮次 | 上下文保留策略 | 所需环境变量 | 推荐用途 |
|---|---|---|---|---|---|
mirothinker_1.7_keep5_max200 ⭐ |
带上下文管理的单代理 | 200 | 保留最近 5 条 | SERPER_API_KEY、SERPER_BASE_URL、JINA_API_KEY、JINA_BASE_URL、E2B_API_KEY、SUMMARY_LLM_BASE_URL、SUMMARY_LLM_MODEL_NAME、SUMMARY_LLM_API_KEY |
1.7(推荐用于大多数任务) |
mirothinker_1.7_keep5_max300 ⭐ |
带上下文管理的单代理 | 300 | 保留最近 5 条 | 同上 | 1.7(用于 BrowseComp 和 BrowseComp-ZH) |
📦 点击展开旧版配置(v0.1/v0.2)
| 配置名称 | 描述 | 最大轮次 | 上下文保留策略 | 所需环境变量 | 推荐用途 |
|---|---|---|---|---|---|
mirothinker_v1.5_keep5_max200 |
带上下文管理的单代理 | 200 | 保留最近 5 条 | SERPER_API_KEY、SERPER_BASE_URL、JINA_API_KEY、JINA_BASE_URL、E2B_API_KEY、SUMMARY_LLM_BASE_URL、SUMMARY_LLM_MODEL_NAME、SUMMARY_LLM_API_KEY |
v1.5(推荐用于大多数任务) |
mirothinker_v1.5_keep5_max400 |
带上下文管理的单代理 | 400 | 保留最近 5 条 | 同上 | v1.5(用于 BrowseComp 和 BrowseComp-ZH) |
mirothinker_v1.5 |
MiroThinker v1.5 单代理 | 600 | 保留所有结果 | 同上 | v1.5 |
mirothinker_v1.0_keep5 |
带上下文管理的单代理 | 600 | 保留最近 5 条 | 同上 | v1.0 |
mirothinker_v1.0 |
MiroThinker v1.0 单代理 | 600 | 保留所有结果 | 同上 | v1.0 |
multi_agent |
使用商业工具的多代理(v0.1/v0.2) | 50 | 保留所有结果 | E2B_API_KEY、ANTHROPIC_API_KEY、ANTHROPIC_BASE_URL、OPENAI_API_KEY、OPENAI_BASE_URL、SERPER_API_KEY、SERPER_BASE_URL、JINA_API_KEY、JINA_BASE_URL |
v0.1/v0.2 |
multi_agent_os |
使用开源工具的多代理(v0.1/v0.2) | 50 | 保留所有结果 | E2B_API_KEY、VISION_API_KEY、VISION_BASE_URL、VISION_MODEL_NAME、WHISPER_API_KEY、WHISPER_BASE_URL、WHISPER_MODEL_NAME、REASONING_API_KEY、REASONING_BASE_URL、REASONING_MODEL_NAME、SERPER_API_KEY、SERPER_BASE_URL、JINA_API_KEY、JINA_BASE_URL |
v0.1/v0.2 |
💡 注意:所有环境变量均列于
apps/miroflow-agent/.env.example中。将其复制到.env并填写您计划使用的工具的值。
创建自定义工具配置
🔧 点击展开自定义工具配置指南
您可以创建自己的 YAML 配置文件,自由组合 MCP 服务器。方法如下:
- 在
apps/miroflow-agent/conf/agent/中创建新的 YAML 文件:
# conf/agent/my_custom_config.yaml
defaults:
- default
- _self_
main_agent:
tools:
- tool-python # Execution environment
- search_and_scrape_webpage # Google search
- jina_scrape_llm_summary # Web scraping with LLM
- tool-vqa # Vision processing (optional)
- tool-transcribe # Audio processing (optional)
- tool-reasoning # Reasoning engine (optional)
- tool-reading # Document reading (optional)
max_turns: 300 # Maximum number of turns
sub_agents:
agent-browsing: # Optional sub-agent
tools:
- tool-google-search
- tool-vqa
- tool-reading
- tool-python
max_turns: 50
keep_tool_result: -1 # Context retention budget: -1 keeps all tool results, or specify K to keep only the K most recent tool responses
💡 上下文保留策略:
keep_tool_result参数实现了一种基于新近度的上下文保留策略。在标准 ReAct 范式中,所有工具输出都会保留在消息历史中,这可能导致上下文利用效率低下。根据我们的实验观察,智能体的后续行动主要依赖于近期的观察结果,而非较早的信息。该策略仅保留最近的 K 条工具响应(其中 K 为keep_tool_result的取值),同时完整保留思考和行动的序列。优势:
- ✅ 保留推理和行动轨迹
- ✅ 引导智能体关注与当前上下文最相关的观察结果
- ✅ 释放额外的上下文空间,以支持更深入的推理和更长的工具使用路径
- ✅ 在不导致性能下降的同时,为交互式扩展提供更多上下文空间
使用方法:设置
keep_tool_result: -1可保留所有工具结果,或指定一个正整数 K(例如keep_tool_result: 5)以仅保留最近的 K 条工具响应。
- 运行评估时使用您的自定义配置:
cd apps/miroflow-agent
uv run main.py llm=qwen-3 agent=my_custom_config llm.base_url=https://your_base_url/v1
-
根据您使用的工具在
.env中配置环境变量。所有可用的环境变量都列在
apps/miroflow-agent/.env.example中。将其复制到.env并根据您选择的配置进行变量配置:cd apps/miroflow-agent cp .env.example .env # 用您实际的 API 密钥编辑 .env对于 MiroThinker v1.5(
mirothinker_v1.5_keep5_max200.yaml、mirothinker_v1.5_keep5_max400.yaml或mirothinker_v1.5.yaml)和 v1.0(mirothinker_v1.0_keep5.yaml或mirothinker_v1.0.yaml),请参见上文的最小配置部分获取完整配置示例。对于其他配置,请参考上文的预配置代理设置表格,查看所需的环境变量。
🔑 点击展开可选 API 密钥
# API for LLM-as-a-Judge (for benchmark testing, required for benchmark evaluation)
OPENAI_API_KEY=your_openai_key
OPENAI_BASE_URL="https://api.openai.com/v1" # Optional, defaults to OpenAI's API
# API for Open-Source Audio Transcription Tool (for benchmark testing, optional)
WHISPER_MODEL_NAME="openai/whisper-large-v3-turbo"
WHISPER_API_KEY=your_whisper_key
WHISPER_BASE_URL="https://your_whisper_base_url/v1"
# API for Open-Source VQA Tool (for benchmark testing, optional)
VISION_MODEL_NAME="Qwen/Qwen2.5-VL-72B-Instruct"
VISION_API_KEY=your_vision_key
VISION_BASE_URL="https://your_vision_base_url/v1/chat/completions"
# API for Open-Source Reasoning Tool (for benchmark testing, optional)
REASONING_MODEL_NAME="Qwen/Qwen3-235B-A22B-Thinking-2507"
REASONING_API_KEY=your_reasoning_key
REASONING_BASE_URL="https://your_reasoning_base_url/v1/chat/completions"
# API for Claude Sonnet 3.7 as Commercial Tools (optional)
ANTHROPIC_API_KEY=your_anthropic_key
# API for Sogou Search (optional)
TENCENTCLOUD_SECRET_ID=your_tencent_cloud_secret_id
TENCENTCLOUD_SECRET_KEY=your_tencent_cloud_secret_key
# API for Summary LLM (can use small models like Qwen3-14B or GPT-5-Nano)
SUMMARY_LLM_BASE_URL="https://your_summary_llm_base_url/v1/chat/completions"
SUMMARY_LLM_MODEL_NAME=your_summary_llm_model_name # e.g., "Qwen/Qwen3-14B" or "gpt-5-nano"
SUMMARY_LLM_API_KEY=your_summary_llm_api_key
部署 MiroThinker 智能体
选项 1(推荐):使用 SGLang 或 vLLM 部署
使用 SGLang 在 61002 端口部署 MiroThinker 模型:
NUM_GPUS=4
PORT=61002
# Downloading agent from HF
AGENT_PATH=miromind-ai/MiroThinker-1.7-mini
python3 -m sglang.launch_server \
--model-path $AGENT_PATH \
--tp $NUM_GPUS \
--dp 1 \
--host 0.0.0.0 \
--port $PORT \
--trust-remote-code
📍 服务器 URL:这将在
http://0.0.0.0:$PORT启动服务器。请将此用作您的服务器基础 URL(例如,http://0.0.0.0:61002/v1)。
选项 2:量化轻量级方案
我们还提供了使用 CPU 优化和 GPU 加速量化技术部署 MiroThinker 智能体的全面指南,以及使用 llama.cpp、Ollama、SGLang 和其他推理框架进行部署的详细分析与操作指引。
📖 完整指南:有关详细部署说明,请参见 部署文档。
运行您的第一个任务
完成环境设置并启动服务器后,运行 main.py 以使用默认问题进行测试:“今天计算机科学领域的 arxiv 论文标题是什么?”
cd apps/miroflow-agent
# Using MiroThinker agents (requires your own server)
uv run python main.py llm=qwen-3 agent=mirothinker_1.7_keep5_max200 llm.base_url=http://localhost:61002/v1
# Or using Claude (requires ANTHROPIC_API_KEY in .env)
uv run python main.py llm=claude-3-7 agent=single_agent_keep5
# Or using GPT-5 (requires OPENAI_API_KEY in .env)
uv run python main.py llm=gpt-5 agent=single_agent_keep5
若要自定义您的问题,请编辑 main.py 的第 32 行:
task_description = "Your custom question here"
该智能体将进行网络搜索,必要时执行代码,并提供带有来源的答案。
📖 更多详情:有关可用配置和故障排除,请参见 apps/miroflow-agent/README.md。
📊 基准测试评估
面向希望复现我们的基准测试结果或在标准基准上进行评估的研究人员。
下载基准测试数据
cd MiroThinker # Back to project root
wget https://huggingface.co/datasets/miromind-ai/MiroFlow-Benchmarks/resolve/main/data_20251115_password_protected.zip
unzip data_20251115_password_protected.zip
# Password: pf4*
rm data_20251115_password_protected.zip
运行基准测试评估
注意: 对于 MiroThinker-1.7,请使用
mirothinker_1.7_keep5_max200(带上下文管理)、mirothinker_1.7_keep5_max300(带上下文管理)。
可用参数:
您可以在运行脚本前通过设置以下环境变量来自定义评估:
| 参数 | 默认值 | 描述 |
|---|---|---|
LLM_MODEL |
"MiroThinker-Agents" |
智能体名称标识符 |
BASE_URL |
"https://your-api.com/v1" |
服务器基础 URL |
NUM_RUNS |
因基准而异 | 评估运行次数(大多数基准为 3 次,GAIA/XBench/FutureX/SEAL-0 为 8 次,AIME2025 为 32 次) |
LLM_PROVIDER |
"qwen" |
大语言模型提供商(例如:qwen、openai、anthropic) |
AGENT_SET |
"mirothinker_1.7_keep5_max200" |
智能体配置(例如:mirothinker_1.7_keep5_max200、mirothinker_1.7_keep5_max300) |
MAX_CONTEXT_LENGTH |
262144 |
最大上下文长度(256K) |
MAX_CONCURRENT |
10 |
最大并发任务数 |
PASS_AT_K |
1 |
Pass@K 评估指标 |
TEMPERATURE |
1.0 |
采样温度 |
API_KEY |
"xxx" |
服务器 API 密钥 |
使用示例:
# Navigate to the miroflow-agent directory first
cd apps/miroflow-agent
# Basic usage with v1.5 (recommended)
NUM_RUNS=8 LLM_MODEL="MiroThinker-1.7-mini" BASE_URL="https://your-api.com/v1" bash scripts/run_evaluate_multiple_runs_gaia-validation-text-103.sh
# Or with v1.0
# NUM_RUNS=8 LLM_MODEL="MiroThinker-v1.0-30B" BASE_URL="https://your-api.com/v1" bash scripts/run_evaluate_multiple_runs_gaia-validation-text-103.sh
# Customize number of runs and agent configuration (v1.5 with context management)
LLM_MODEL="MiroThinker-1.7-mini" \
BASE_URL="https://your-api.com/v1" \
NUM_RUNS=8 \
AGENT_SET="mirothinker_1.7_keep5_max200" \
bash scripts/run_evaluate_multiple_runs_gaia-validation-text-103.sh
📋 点击展开所有基准测试命令
⚠️ 对 MiroThinker-1.7 的重要提示:要复现我们报告的结果,必须设置正确的
AGENT_SET:
- BrowseComp 和 BrowseComp-ZH:使用
AGENT_SET="mirothinker_1.7_keep5_max300"- 所有其他基准测试:使用
AGENT_SET="mirothinker_1.7_keep5_max200"
# Navigate to the miroflow-agent directory first
cd apps/miroflow-agent
# HLE
NUM_RUNS=3 LLM_MODEL="xxx" BASE_URL="xxx" AGENT_SET="mirothinker_1.7_keep5_max200" bash scripts/run_evaluate_multiple_runs_hle.sh
# HLE-Text-2158
NUM_RUNS=3 LLM_MODEL="xxx" BASE_URL="xxx" AGENT_SET="mirothinker_1.7_keep5_max200" bash scripts/run_evaluate_multiple_runs_hle-text-2158.sh
# HLE-Text-500
NUM_RUNS=3 LLM_MODEL="xxx" BASE_URL="xxx" AGENT_SET="mirothinker_1.7_keep5_max200" bash scripts/run_evaluate_multiple_runs_hle-text-500.sh
# GAIA-Text-103
NUM_RUNS=8 LLM_MODEL="xxx" BASE_URL="xxx" AGENT_SET="mirothinker_1.7_keep5_max200" bash scripts/run_evaluate_multiple_runs_gaia-validation-text-103.sh
# GAIA-Validation (GAIA-Val-165)
NUM_RUNS=8 LLM_MODEL="xxx" BASE_URL="xxx" AGENT_SET="mirothinker_1.7_keep5_max200" bash scripts/run_evaluate_multiple_runs_gaia-validation.sh
# BrowseComp-EN (⚠️ use max300)
NUM_RUNS=3 LLM_MODEL="xxx" BASE_URL="xxx" AGENT_SET="mirothinker_1.7_keep5_max300" bash scripts/run_evaluate_multiple_runs_browsecomp.sh
# BrowseComp-ZH (⚠️ use max300)
NUM_RUNS=3 LLM_MODEL="xxx" BASE_URL="xxx" AGENT_SET="mirothinker_1.7_keep5_max300" bash scripts/run_evaluate_multiple_runs_browsecomp_zh.sh
# WebWalkerQA
NUM_RUNS=3 LLM_MODEL="xxx" BASE_URL="xxx" AGENT_SET="mirothinker_1.7_keep5_max200" bash scripts/run_evaluate_multiple_runs_webwalkerqa.sh
# XBench-DeepSearch
NUM_RUNS=8 LLM_MODEL="xxx" BASE_URL="xxx" AGENT_SET="mirothinker_1.7_keep5_max200" bash scripts/run_evaluate_multiple_runs_xbench_deepsearch.sh
# FRAMES
NUM_RUNS=3 LLM_MODEL="xxx" BASE_URL="xxx" AGENT_SET="mirothinker_1.7_keep5_max200" bash scripts/run_evaluate_multiple_runs_frames.sh
# SEAL-0
NUM_RUNS=8 LLM_MODEL="xxx" BASE_URL="xxx" AGENT_SET="mirothinker_1.7_keep5_max200" bash scripts/run_evaluate_multiple_runs_seal-0.sh
# FutureX
NUM_RUNS=8 LLM_MODEL="xxx" BASE_URL="xxx" AGENT_SET="mirothinker_1.7_keep5_max200" bash scripts/run_evaluate_multiple_runs_futurex.sh
# AIME2025
NUM_RUNS=32 LLM_MODEL="xxx" BASE_URL="xxx" AGENT_SET="mirothinker_1.7_keep5_max200" bash scripts/run_evaluate_multiple_runs_aime2025.sh
# DeepSearchQA
NUM_RUNS=3 LLM_MODEL="xxx" BASE_URL="xxx" AGENT_SET="mirothinker_1.7_keep5_max200" bash scripts/run_evaluate_multiple_runs_deepsearchqa.sh
3. 监控评估进度
📊 点击展开进度监控命令
# Navigate to the miroflow-agent directory first
cd apps/miroflow-agent
# For HLE
python benchmarks/check_progress/check_progress_hle.py /path/to/evaluation/logs
# For HLE-Text-2158
python benchmarks/check_progress/check_progress_hle-text-2158.py /path/to/evaluation/logs
# For HLE-Text-500
python benchmarks/check_progress/check_progress_hle-text-500.py /path/to/evaluation/logs
# For BrowseComp-EN
python benchmarks/check_progress/check_progress_browsecomp.py /path/to/evaluation/logs
# For BrowseComp-ZH
python benchmarks/check_progress/check_progress_browsecomp_zh.py /path/to/evaluation/logs
# For GAIA-Validation
python benchmarks/check_progress/check_progress_gaia-validation.py /path/to/evaluation/logs
# For GAIA-Text-103
python benchmarks/check_progress/check_progress_gaia-validation-text-103.py /path/to/evaluation/logs
# For WebWalkerQA
python benchmarks/check_progress/check_progress_webwalkerqa.py /path/to/evaluation/logs
# For Frames
python benchmarks/check_progress/check_progress_frames.py /path/to/evaluation/logs
# For XBench-DeepSearch
python benchmarks/check_progress/check_progress_xbench_deepsearch.py /path/to/evaluation/logs
# For SEAL-0
python benchmarks/check_progress/check_progress_seal-0.py /path/to/evaluation/logs
# For AIME2025
python benchmarks/check_progress/check_progress_aime2025.py /path/to/evaluation/logs
# For DeepSearchQA
python benchmarks/check_progress/check_progress_deepsearchqa.py /path/to/evaluation/logs
🔬 轨迹收集
📋 点击展开轨迹收集命令
cd apps/collect-trace
# Collect Traces for SFT
bash scripts/collect_trace_claude37.sh
bash scripts/collect_trace_gpt5.sh
# Collect Traces for DPO
bash scripts/collect_trace_qwen3.sh
❓ 常见问题与故障排除
常见问题
🔧 点击展开故障排除指南
问:我应该使用哪个版本?
答: 我们推荐使用 MiroThinker-1.7 ⭐,并搭配以下最低配置:
- v1.7 ⭐:最新版本,支持256K上下文,具备业界领先性能。使用以下配置(含上下文管理):
mirothinker_1.7_keep5_max200(最多200轮对话,适用于大多数任务)mirothinker_1.7_keep5_max300(最多300轮对话,仅用于BrowseComp和BrowseComp-ZH)
问:如何获取API密钥?
答: 您需要以下密钥以完成基础设置:
- SERPER_API_KEY:从 Serper.dev 获取(谷歌搜索API)
- JINA_API_KEY:从 Jina.ai 获取(网页抓取)
- E2B_API_KEY:从 E2B.dev 获取(代码执行沙箱)
- SUMMARY_LLM_API_KEY:您的LLM API凭证(用于内容总结)。可使用小型模型如Qwen3-14B或GPT-5-Nano,模型选择对性能影响极小。
- OPENAI_API_KEY:从 OpenAI 获取(基准测试评估必需,用于LLM作为评判工具)
- OPENAI_BASE_URL:可选,默认值为
https://api.openai.com/v1。可更改为使用兼容OpenAI的API。
问:代理服务器连接错误
答: 常见问题:
- 检查基础URL格式:应以
/v1结尾(例如https://your-api.com/v1) - 验证API密钥:确保在环境变量或脚本中正确设置
API_KEY - 检查服务器状态:确保您的服务器正在运行且可访问
- 网络问题:确认防火墙/网络设置允许连接
问:评估脚本运行失败
答: 故障排除步骤:
- 检查工作目录:确保您位于
apps/miroflow-agent目录下 - 验证环境:运行
uv sync确保依赖项已安装 - 检查.env文件:确保所有必需的环境变量均已设置
- 查看日志:检查
logs/目录获取详细错误信息 - 验证数据路径:确保基准测试数据已下载并位于正确位置
问:在WSL上运行 uv sync 时出现内存分配错误
答: WSL2设置了内存上限,在构建大型包(如 transformers、pillow)时可能导致 uv sync 失败。有两种解决方法:
-
增加WSL2内存限制(推荐):在Windows主机上创建或编辑
%UserProfile%\.wslconfig,然后重启WSL(wsl --shutdown):[wsl2] memory=8GB -
限制并行包构建(无需重启):在运行
uv sync前设置UV_CONCURRENT_BUILDS环境变量:UV_CONCURRENT_BUILDS=1 uv sync
问:内存不足错误
答: 解决方法:
- 减少上下文长度:将
MAX_CONTEXT_LENGTH设置为较小值(例如131072表示128K) - 使用轮次更少的上下文管理:
- 对于v1.5:使用
mirothinker_1.7_keep5_max200或mirothinker_1.7_keep5_max300(带上下文管理)
- 对于v1.5:使用
- 减少并发任务:将
MAX_CONCURRENT设置为较小数值(例如5) - 使用更小的代理模型:
- 对于v1.5:尝试30B版本而非235B
- 对于v1.0:尝试8B或30B版本而非72B
问:工具执行错误
答: 常见修复方法:
- E2B错误:验证
E2B_API_KEY有效且账户有余额 - Serper错误:检查
SERPER_API_KEY和速率限制 - Jina错误:验证
JINA_API_KEY和JINA_BASE_URL正确 - LLM总结错误:检查
SUMMARY_LLM_*变量和代理可用性
问:如何监控长时间运行的评估?
答: 使用进度监控脚本:
cd apps/miroflow-agent
python benchmarks/check_progress/check_progress_<benchmark_name>.py /path/to/logs
脚本会显示完成状态、已用时间和预计剩余时间。
获取帮助
- 📖 文档:查看 MiroFlow Tools README 了解工具详情
- 💬 Discord:加入我们的 Discord 社区
- 🐛 问题反馈:在 GitHub Issues 上报告漏洞
- 📧 联系我们:访问 官方网站 获取更多信息
📄 许可证
本项目采用 Apache 2.0 许可证授权 - 详情参见 LICENSE 文件。
🙏 致谢
我们向以下各方致以诚挚的感谢:
- 🏆 基准测试贡献者 提供的全面评估数据集
- 🌍 开源社区 提供的工具和库,使本项目得以实现
- 👥 所有贡献者 为改进 MiroThinker 所付出的努力
加入我们的社区,共同打造 AI 智能体的未来!
参考文献
如果您的研究中使用了本项目,请考虑引用:
MiroThinker(模型与方法)
@article{miromind2026mirothinker,
title={MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification},
author={MiroMind Team and Bai, S. and Bing, L. and Lei, L. and Li, R. and Li, X. and Lin, X. and Min, E. and Su, L. and Wang, B. and Wang, L. and Wang, L. and Wang, S. and Wang, X. and Zhang, Y. and Zhang, Z. and others},
journal={arXiv preprint arXiv:2603.15726},
year={2026}
}
@article{miromind2025mirothinker,
title={MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling},
author={MiroMind Team and Bai, Song and Bing, Lidong and Chen, Carson and Chen, Guanzheng and Chen, Yuntao and Chen, Zhe and Chen, Ziyi and Dong, Xuan and others},
journal={arXiv preprint arXiv:2511.11793},
year={2025}
}
MiroFlow(框架)
@article{miromind2026miroflow,
title={MiroFlow: Towards High-Performance and Robust Open-Source Agent Framework for General Deep Research Tasks},
author={Su, Shiqian and Xing, Sen and Dong, Xuan and Zhong, Muyan and Wang, Bin and Zhu, Xizhou and Chen, Yuntao and Wang, Wenhai and Deng, Yue and Zhu, Pengxiang and others},
journal={arXiv preprint arXiv:2602.22808},
year={2026}
}
Introduction
MiroThinker 是一款开源智能体模型,专为深度研究和复杂工具使用场景训练而成。【此简介由AI生成】