outlines:基于 Python 类型系统的大语言模型结构化输出项目

Structured Outputs

分支13Tags89
文件最后提交记录最后更新时间
1 年前
2 个月前
1 年前
1 个月前
1 年前
1 个月前
1 个月前
1 个月前
1 年前
1 个月前
1 年前
1 年前
3 年前
1 个月前
2 年前
4 个月前
1 年前
1 年前
1 年前
1 年前
1 个月前
1 个月前
1 年前
3 年前
1 年前
1 个月前

Outlines Logo Outlines Logo

🗒️ 大语言模型的结构化输出 🗒️

.txt 团队用心打造 👷️
受到 NVIDIA、Cohere、HuggingFace、vLLM 等企业的信赖

[![PyPI 版本][pypi-version-badge]][pypi] [![下载量][downloads-badge]][pypistats] [![星标数][stars-badge]][stars]

[![Discord][discord-badge]][discord] [![博客][dottxt-blog-badge]][dottxt-blog] [![Twitter][twitter-badge]][twitter]


.txt API 目前处于早期访问阶段。立即申请访问 →

🚀 构建结构化生成的未来

我们正与精选合作伙伴携手,开发结构化生成的全新接口。

需要 XML、FHIR、自定义模式或语法?欢迎交流。

模式审计:分享您的模式,我们将为您展示生成过程中可能出现的问题、修复这些问题的约束条件,以及优化前后的合规率。立即注册

目录

为什么选择 Outlines?

大型语言模型(LLMs)功能强大,但其输出结果难以预测。大多数解决方案尝试在生成后通过解析、正则表达式或脆弱且易出错的代码来修复不良输出。

Outlines 能够在生成过程中直接从任何大型语言模型保证结构化输出。

  • 适用于任何模型 - 相同代码可在 OpenAI、Ollama、vLLM 等平台运行
  • 简单集成 - 只需传递所需的输出类型:model(prompt, output_type)
  • 保证结构有效 - 不再有解析难题或损坏的 JSON
  • 独立于服务提供商 - 无需更改代码即可切换模型

Outlines 的理念

Outlines 遵循一种简单的模式,该模式借鉴了 Python 自身的类型系统。只需指定所需的输出类型,Outlines 将确保您的数据完全匹配该结构:

  • 对于是/否响应,使用 Literal["Yes", "No"]
  • 对于数值,使用 int
  • 对于复杂对象,使用 Pydantic 模型 定义结构

快速入门

使用 Outlines 入门非常简单:

1. 安装 outlines

pip install outlines

2. 连接到您偏好的模型

import outlines
from transformers import AutoTokenizer, AutoModelForCausalLM


MODEL_NAME = "microsoft/Phi-3-mini-4k-instruct"
model = outlines.from_transformers(
    AutoModelForCausalLM.from_pretrained(MODEL_NAME, device_map="auto"),
    AutoTokenizer.from_pretrained(MODEL_NAME)
)

3. 从简单的结构化输出开始

from typing import Literal
from pydantic import BaseModel


# Simple classification
sentiment = model(
    "Analyze: 'This product completely changed my life!'",
    Literal["Positive", "Negative", "Neutral"]
)
print(sentiment)  # "Positive"

# Extract specific types
temperature = model("What's the boiling point of water in Celsius?", int)
print(temperature)  # 100

4. 创建复杂结构

from pydantic import BaseModel
from enum import Enum

class Rating(Enum):
    poor = 1
    fair = 2
    good = 3
    excellent = 4

class ProductReview(BaseModel):
    rating: Rating
    pros: list[str]
    cons: list[str]
    summary: str

review = model(
    "Review: The XPS 13 has great battery life and a stunning display, but it runs hot and the webcam is poor quality.",
    ProductReview,
    max_new_tokens=200,
)

review = ProductReview.model_validate_json(review)
print(f"Rating: {review.rating.name}")  # "Rating: good"
print(f"Pros: {review.pros}")           # "Pros: ['great battery life', 'stunning display']"
print(f"Summary: {review.summary}")     # "Summary: Good laptop with great display but thermal issues"

实际应用示例

以下是可直接用于生产环境的示例,展示 Outlines 如何解决常见问题:

🙋‍♂️ 客户支持分类
本示例演示如何将自由格式的客户邮件转换为结构化服务工单。通过解析优先级、类别和升级标志等属性,该代码可实现支持问题的自动路由和处理。
import outlines
from enum import Enum
from pydantic import BaseModel
from transformers import AutoTokenizer, AutoModelForCausalLM
from typing import List


MODEL_NAME = "microsoft/Phi-3-mini-4k-instruct"
model = outlines.from_transformers(
    AutoModelForCausalLM.from_pretrained(MODEL_NAME, device_map="auto"),
    AutoTokenizer.from_pretrained(MODEL_NAME)
)


def alert_manager(ticket):
    print("Alert!", ticket)


class TicketPriority(str, Enum):
    low = "low"
    medium = "medium"
    high = "high"
    urgent = "urgent"

class ServiceTicket(BaseModel):
    priority: TicketPriority
    category: str
    requires_manager: bool
    summary: str
    action_items: List[str]


customer_email = """
Subject: URGENT - Cannot access my account after payment

I paid for the premium plan 3 hours ago and still can't access any features.
I've tried logging out and back in multiple times. This is unacceptable as I
have a client presentation in an hour and need the analytics dashboard.
Please fix this immediately or refund my payment.
"""

prompt = f"""
<|im_start|>user
Analyze this customer email:

{customer_email}
<|im_end|>
<|im_start|>assistant
"""

ticket = model(
    prompt,
    ServiceTicket,
    max_new_tokens=500
)

# Use structured data to route the ticket
ticket = ServiceTicket.model_validate_json(ticket)
if ticket.priority == "urgent" or ticket.requires_manager:
    alert_manager(ticket)
📦 电商产品分类
本用例展示了 outlines 如何将产品描述转换为结构化分类数据(例如主类别、子类别和属性),以简化库存管理等任务。每个产品描述都会被自动处理,减少了手动分类的工作量。
import outlines
from pydantic import BaseModel
from transformers import AutoTokenizer, AutoModelForCausalLM
from typing import List, Optional


MODEL_NAME = "microsoft/Phi-3-mini-4k-instruct"
model = outlines.from_transformers(
    AutoModelForCausalLM.from_pretrained(MODEL_NAME, device_map="auto"),
    AutoTokenizer.from_pretrained(MODEL_NAME)
)


def update_inventory(product, category, sub_category):
    print(f"Updated {product.split(',')[0]} in category {category}/{sub_category}")


class ProductCategory(BaseModel):
    main_category: str
    sub_category: str
    attributes: List[str]
    brand_match: Optional[str]

# Process product descriptions in batches
product_descriptions = [
    "Apple iPhone 15 Pro Max 256GB Titanium, 6.7-inch Super Retina XDR display with ProMotion",
    "Organic Cotton T-Shirt, Men's Medium, Navy Blue, 100% Sustainable Materials",
    "KitchenAid Stand Mixer, 5 Quart, Red, 10-Speed Settings with Dough Hook Attachment"
]

template = outlines.Template.from_string("""
<|im_start|>user
Categorize this product:

{{ description }}
<|im_end|>
<|im_start|>assistant
""")

# Get structured categorization for all products
categories = model(
    [template(description=desc) for desc in product_descriptions],
    ProductCategory,
    max_new_tokens=200
)

# Use categorization for inventory management
categories = [
    ProductCategory.model_validate_json(category) for category in categories
]
for product, category in zip(product_descriptions, categories):
    update_inventory(product, category.main_category, category.sub_category)
📊 解析数据不完整的事件详情
本示例使用 outlines 将事件描述解析为结构化信息(如事件名称、日期、地点、类型和主题),即使在数据不完整的情况下也能处理。它利用联合类型返回结构化事件数据或备用的“我不知道”答案,确保在不同场景下都能可靠地提取信息。
import outlines
from typing import Union, List, Literal
from pydantic import BaseModel
from enum import Enum
from transformers import AutoTokenizer, AutoModelForCausalLM


MODEL_NAME = "microsoft/Phi-3-mini-4k-instruct"
model = outlines.from_transformers(
    AutoModelForCausalLM.from_pretrained(MODEL_NAME, device_map="auto"),
    AutoTokenizer.from_pretrained(MODEL_NAME)
)

class EventType(str, Enum):
    conference = "conference"
    webinar = "webinar"
    workshop = "workshop"
    meetup = "meetup"
    other = "other"


class EventInfo(BaseModel):
    """Structured information about a tech event"""
    name: str
    date: str
    location: str
    event_type: EventType
    topics: List[str]
    registration_required: bool

# Create a union type that can either be a structured EventInfo or "I don't know"
EventResponse = Union[EventInfo, Literal["I don't know"]]

# Sample event descriptions
event_descriptions = [
    # Complete information
    """
    Join us for DevCon 2023, the premier developer conference happening on November 15-17, 2023
    at the San Francisco Convention Center. Topics include AI/ML, cloud infrastructure, and web3.
    Registration is required.
    """,

    # Insufficient information
    """
    Tech event next week. More details coming soon!
    """
]

# Process events
results = []
for description in event_descriptions:
    prompt = f"""
<|im_start>system
You are a helpful assistant
<|im_end|>
<|im_start>user
Extract structured information about this tech event:

{description}

If there is enough information, return a JSON object with the following fields:

- name: The name of the event
- date: The date where the event is taking place
- location: Where the event is taking place
- event_type: either 'conference', 'webinar', 'workshop', 'meetup' or 'other'
- topics: a list of topics of the conference
- registration_required: a boolean that indicates whether registration is required

If the information available does not allow you to fill this JSON, and only then, answer 'I don't know'.
<|im_end|>
<|im_start|>assistant
"""
    # Union type allows the model to return structured data or "I don't know"
    result = model(prompt, EventResponse, max_new_tokens=200)
    results.append(result)

# Display results
for i, result in enumerate(results):
    print(f"Event {i+1}:")
    if isinstance(result, str):
        print(f"  {result}")
    else:
        # It's an EventInfo object
        print(f"  Name: {result.name}")
        print(f"  Type: {result.event_type}")
        print(f"  Date: {result.date}")
        print(f"  Topics: {', '.join(result.topics)}")
    print()

# Use structured data in downstream processing
structured_count = sum(1 for r in results if isinstance(r, EventInfo))
print(f"Successfully extracted data for {structured_count} of {len(results)} events")
🗂️ 将文档分类为预定义类型
在这种情况下,outlines 使用文字类型规范将文档分类为预定义类别(例如,“财务报告”、“法律合同”)。分类结果以表格格式和类别分布摘要的形式展示,说明了结构化输出如何简化内容管理。
import outlines
from typing import Literal, List
import pandas as pd
from transformers import AutoTokenizer, AutoModelForCausalLM


MODEL_NAME = "microsoft/Phi-3-mini-4k-instruct"
model = outlines.from_transformers(
    AutoModelForCausalLM.from_pretrained(MODEL_NAME, device_map="auto"),
    AutoTokenizer.from_pretrained(MODEL_NAME)
)


# Define classification categories using Literal
DocumentCategory = Literal[
    "Financial Report",
    "Legal Contract",
    "Technical Documentation",
    "Marketing Material",
    "Personal Correspondence"
]

# Sample documents to classify
documents = [
    "Q3 Financial Summary: Revenue increased by 15% year-over-year to $12.4M. EBITDA margin improved to 23% compared to 19% in Q3 last year. Operating expenses...",

    "This agreement is made between Party A and Party B, hereinafter referred to as 'the Parties', on this day of...",

    "The API accepts POST requests with JSON payloads. Required parameters include 'user_id' and 'transaction_type'. The endpoint returns a 200 status code on success."
]

template = outlines.Template.from_string("""
<|im_start|>user
Classify the following document into exactly one category among the following categories:
- Financial Report
- Legal Contract
- Technical Documentation
- Marketing Material
- Personal Correspondence

Document:
{{ document }}
<|im_end|>
<|im_start|>assistant
""")

# Classify documents
def classify_documents(texts: List[str]) -> List[DocumentCategory]:
    results = []

    for text in texts:
        prompt = template(document=text)
        # The model must return one of the predefined categories
        category = model(prompt, DocumentCategory, max_new_tokens=200)
        results.append(category)

    return results

# Perform classification
classifications = classify_documents(documents)

# Create a simple results table
results_df = pd.DataFrame({
    "Document": [doc[:50] + "..." for doc in documents],
    "Classification": classifications
})

print(results_df)

# Count documents by category
category_counts = pd.Series(classifications).value_counts()
print("\nCategory Distribution:")
print(category_counts)
📅 通过函数调用根据请求安排会议
本示例展示了 outlines 如何解析自然语言会议请求,并将其转换为与预定义函数参数匹配的结构化格式。提取会议详情(如标题、日期、时长、参会人员)后,这些信息将用于自动安排会议。
import outlines
import json
from typing import List, Optional
from datetime import date
from transformers import AutoTokenizer, AutoModelForCausalLM


MODEL_NAME = "microsoft/phi-4"
model = outlines.from_transformers(
    AutoModelForCausalLM.from_pretrained(MODEL_NAME, device_map="auto"),
    AutoTokenizer.from_pretrained(MODEL_NAME)
)


# Define a function with typed parameters
def schedule_meeting(
    title: str,
    date: date,
    duration_minutes: int,
    attendees: List[str],
    location: Optional[str] = None,
    agenda_items: Optional[List[str]] = None
):
    """Schedule a meeting with the specified details"""
    # In a real app, this would create the meeting
    meeting = {
        "title": title,
        "date": date,
        "duration_minutes": duration_minutes,
        "attendees": attendees,
        "location": location,
        "agenda_items": agenda_items
    }
    return f"Meeting '{title}' scheduled for {date} with {len(attendees)} attendees"

# Natural language request
user_request = """
I need to set up a product roadmap review with the engineering team for next
Tuesday at 2pm. It should last 90 minutes. Please invite john@example.com,
sarah@example.com, and the product team at product@example.com.
"""

# Outlines automatically infers the required structure from the function signature
prompt = f"""
<|im_start|>user
Extract the meeting details from this request:

{user_request}
<|im_end|>
<|im_start|>assistant
"""
meeting_params = model(prompt, schedule_meeting, max_new_tokens=200)

# The result is a dictionary matching the function parameters
meeting_params = json.loads(meeting_params)
print(meeting_params)

# Call the function with the extracted parameters
result = schedule_meeting(**meeting_params)
print(result)
# "Meeting 'Product Roadmap Review' scheduled for 2023-10-17 with 3 attendees"
📝 使用可复用模板动态生成提示词
本示例使用基于 Jinja 的模板,展示了如何为情感分析等任务生成动态提示词。它说明了如何轻松复用和自定义提示词(包括小样本学习策略)以适应不同的内容类型,同时确保输出保持结构化。
import outlines
from typing import List, Literal
from transformers import AutoTokenizer, AutoModelForCausalLM


MODEL_NAME = "microsoft/phi-4"
model = outlines.from_transformers(
    AutoModelForCausalLM.from_pretrained(MODEL_NAME, device_map="auto"),
    AutoTokenizer.from_pretrained(MODEL_NAME)
)


# 1. Create a reusable template with Jinja syntax
sentiment_template = outlines.Template.from_string("""
<|im_start>user
Analyze the sentiment of the following {{ content_type }}:

{{ text }}

Provide your analysis as either "Positive", "Negative", or "Neutral".
<|im_end>
<|im_start>assistant
""")

# 2. Generate prompts with different parameters
review = "This restaurant exceeded all my expectations. Fantastic service!"
prompt = sentiment_template(content_type="review", text=review)

# 3. Use the templated prompt with structured generation
result = model(prompt, Literal["Positive", "Negative", "Neutral"])
print(result)  # "Positive"

# Templates can also be loaded from files
example_template = outlines.Template.from_file("templates/few_shot.txt")

# Use with examples for few-shot learning
examples = [
    ("The food was cold", "Negative"),
    ("The staff was friendly", "Positive")
]
few_shot_prompt = example_template(examples=examples, query="Service was slow")
print(few_shot_prompt)

他们都在用 outlines

Users Logo Users Logo

模型集成

模型类型 描述 文档
服务器支持 vLLM 和 Ollama 服务器集成 →
本地模型支持 transformers 和 llama.cpp 模型集成 →
API 支持 OpenAI、Gemini 和 Dottxt API 集成 →

核心功能

功能 描述 文档
多项选择 将输出限制为预定义选项 多项选择指南 →
函数调用 从函数签名推断结构 函数指南 →
JSON/Pydantic 生成符合 JSON 模式的输出 JSON 指南 →
正则表达式 生成遵循正则表达式模式的文本 正则表达式指南 →
语法规则 强制复杂的输出结构 语法规则指南 →

其他功能

功能 描述 文档
提示词模板 将复杂提示词与代码分离 模板指南 →
自定义类型 构建复杂类型的直观界面 Python 类型指南 →
应用程序 将模板和类型封装到函数中 应用程序指南 →

关于 .txt

dottxt logo dottxt logo

Outlines 由 .txt 开发并维护,这是一家致力于提升大型语言模型(LLMs)在生产应用中可靠性的公司。

我们专注于通过以下方式推动结构化生成技术的发展:

关注我们的 Twitter 或查看我们的 blog,以获取我们在提升 LLMs 可靠性方面的最新工作动态。

社区

[![Contributors][contributors-badge]][contributors] [![Stars][stars-badge]][stars] [![Downloads][downloads-badge]][pypistats] [![Discord badge][discord-badge]][discord]

  • 💡 有想法? 来 [Discord][discord] 与我们交流
  • 🐞 发现漏洞? 提交 issue
  • 🧩 想要贡献? 参考我们的 contribution guide

引用 Outlines

@article{willard2023efficient,
  title={Efficient Guided Generation for Large Language Models},
  author={Willard, Brandon T and Louf, R{\'e}mi},
  journal={arXiv preprint arXiv:2307.09702},
  year={2023}
}

[贡献者]:https://github.com/dottxt-ai/outlines/graphs/contributors [贡献者徽章]:https://img.shields.io/github/contributors/dottxt-ai/outlines?style=flat-square&logo=github&logoColor=white&color=ECEFF4 [dottxt博客]:https://blog.dottxt.co/ [dottxt博客徽章]:https://img.shields.io/badge/dottxt blog-a6b4a3 [dottxt推特]:https://twitter.com/dottxtai [dottxt推特徽章]:https://img.shields.io/twitter/follow/dottxtai?style=social [Discord]:https://discord.gg/R9DSu34mGd [Discord徽章]:https://img.shields.io/discord/1182316225284554793?color=ddb8ca&logo=discord&logoColor=white&style=flat-square [下载量徽章]:https://img.shields.io/pypi/dm/outlines?color=A6B4A3&logo=python&logoColor=white&style=flat-square [PyPI统计]:https://pypistats.org/packages/outlines [PyPI版本徽章]:https://img.shields.io/pypi/v/outlines?style=flat-square&logoColor=white&color=ddb8ca [PyPI]:https://pypi.org/project/outlines/ [星标数]:https://github.com/dottxt-ai/outlines/stargazers [星标数徽章]:https://img.shields.io/github/stars/dottxt-ai/outlines?style=flat-square&logo=github&color=BD932F&logoColor=white [推特徽章]:https://img.shields.io/twitter/follow/dottxtai?style=flat-square&logo=x&logoColor=white&color=bd932f [推特]:https://x.com/dottxtai