| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
feat: Implement retry-after mechanism for model apis and udfs (#5769) ## Changes Made Adds a RetryAfterException that can be raised from a UDF, which will supply the UDF retry mechanism with a delay to sleep. This allows us to respect the retry-afters provided in headers from 429 and 503 errors from clients, such as our openai and google clients in the model apis. By default, all of our model apis should already be configured with max_retries = 3, and our exponential backoff has an initial delay of 100ms. It might be worth exploring in the future how we may want to more easily expose these configurations to users. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
feat: Add configurable token limits to OpenAI text embedder (#6017) ## Changes Made Add batch_token_limit to EmbedTextOptions and input_text_token_limit to _ModelProfile to support different providers and models with varying token constraints. Defaults remain 300,000 and 8,192 respectively. This allows LM Studio, Together AI, and other OpenAI-compatible services to configure appropriate limits for their models and rate limits. ## Related Issues Closes #5990 <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 8 个月前 | |
feat: Add configurable token limits to OpenAI text embedder (#6017) ## Changes Made Add batch_token_limit to EmbedTextOptions and input_text_token_limit to _ModelProfile to support different providers and models with varying token constraints. Defaults remain 300,000 and 8,192 respectively. This allows LM Studio, Together AI, and other OpenAI-compatible services to configure appropriate limits for their models and rate limits. ## Related Issues Closes #5990 <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 8 个月前 | |
fix(ai): resolve intermittent meta tensor error in classify_text/classify_image (#5977) ## Summary Fixes #5707 Resolves intermittent NotImplementedError: Cannot copy out of meta tensor; no data! error in classify_text() and classify_image() when dynamic batching is enabled. ## Problem classify_text() and classify_image() fail intermittently (~20-40% of the time) with: NotImplementedError: Cannot copy out of meta tensor; no data! **Key characteristics:** - Intermittent failure (not deterministic) - More reproducible with dynamic batching enabled - Fails during model initialization, not inference - Particularly affects BART models (e.g., facebook/bart-large-mnli) ## Root Cause When Daft's dynamic batching creates multiple UDF worker instances concurrently, they all try to load the transformers pipeline simultaneously. This triggers a **race condition in transformers' model loading code** where some model parameters get stuck as meta tensors (placeholders without actual data) instead of being properly materialized. ## Solution Add a global lock to serialize model loading across all worker threads: python import threading _model_loading_lock = threading.Lock() def __init__(self, model_name_or_path: str, **options): with _model_loading_lock: self._pipeline = pipeline( task="zero-shot-classification", model=model_name_or_path, device=get_torch_device(), ) **Applied to:** - daft/ai/transformers/protocols/text_classifier.py - daft/ai/transformers/protocols/image_classifier.py ## Testing Created synthetic test with dynamic batching enabled: python daft.set_execution_config( default_morsel_size=12, enable_dynamic_batching=True, dynamic_batching_strategy="auto" ) **Results:** - **Without lock:** ~40% failure rate (30 iterations) - **With lock:** 0% failure rate (30 iterations) ✓ ## Investigation Details ### Test Results Summary Tested various approaches before arriving at the lock solution: - device_map=None only: ~40% failure rate - device_map=None + low_cpu_mem_usage=False: ~43% failure rate - device_map=None + low_cpu_mem_usage=False + CPU-first loading: ~23% failure rate - **Lock only: 0% failure rate** ← Final solution ### Why This Works The lock ensures only one worker loads the model at a time, preventing the concurrent access that triggers the meta tensor race condition in transformers. Each worker still loads its own model copy (no sharing), but initialization happens serially. ### Performance Impact **Negligible:** - Lock only affects initialization time (one-time per worker) - Inference runs in parallel without locks - Workers typically initialize serially anyway due to memory constraints ## Caveats **Note:** This may be papering over a deeper issue in transformers. The lock eliminates the race condition but doesn't address the underlying meta tensor bug in the upstream library. An upstream fix in transformers would be ideal long-term. ## References - **Issue #5707:** https://github.com/Eventual-Inc/Daft/issues/5707 - **Transformers issue #36247:** https://github.com/huggingface/transformers/issues/36247 - **Environment:** Transformers 4.57.1, PyTorch 2.8.0, MPS (Apple Silicon) | 8 个月前 | |
feat: Improve model api typing (#5809) ## Changes Made Types the **options param in the model APIs using TypedDicts. Each model api now has it's own specific TypedDict that indicates the a subset of the possible options that can be passed, providing better typing hints for the user and their IDEs. For most of the model apis, the options include max_retries, on_error, and batch_size, which are subsequently passed into the UDF. Provider specific options can be subclassed, for example the OpenAIPromptOptions is a subclass of PromptOptions, and has an additional field for use_chat_completions, which is specific to openai prompting. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
feat: model resource plumbing for ml model-based functions (#4902) ## Changes Made * Adds embed_text daft expression * Adds TextEmbedder protocol * Adds Provider base class * Adds Descriptor[T] for making deferred initialization an API invariant * Removed batched vs. non-batched protocols to simplify * Removed lifecycle methods from the protocols **SentenceTransformers** * Adds SentenceTransformersTexEmbedder implementation * Adds SentenceTransformersProvider implementation Will implement for additional providers upon review. * Unit tests * Add OpenAI provider support * Add VLLM provider support ## Related Issues n/a ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly (tag @/ccmao1130 for docs review) | 1 年前 | |
| 8 个月前 | ||
feat: embed text metrics (#5583) ## Changes Made Metrics for embed_text ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly | 10 个月前 | |
| 8 个月前 | ||
chore: bump mypy and ruff in pre-commit (#5836) ## Changes Made Bump mypy to 1.19.1 and ruff to 0.14.10 and resolve pre-commit errors. Rationale: my local ruff version is much higher than that in the project pre-commit, so I am getting a lot of linter warnings in my IDE. It is also about time to upgrade lints to utilize newer Python language features. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
feat: Add configurable token limits to OpenAI text embedder (#6017) ## Changes Made Add batch_token_limit to EmbedTextOptions and input_text_token_limit to _ModelProfile to support different providers and models with varying token constraints. Defaults remain 300,000 and 8,192 respectively. This allows LM Studio, Together AI, and other OpenAI-compatible services to configure appropriate limits for their models and rate limits. ## Related Issues Closes #5990 <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 8 个月前 | |
feat: Implement retry-after mechanism for model apis and udfs (#5769) ## Changes Made Adds a RetryAfterException that can be raised from a UDF, which will supply the UDF retry mechanism with a delay to sleep. This allows us to respect the retry-afters provided in headers from 429 and 503 errors from clients, such as our openai and google clients in the model apis. By default, all of our model apis should already be configured with max_retries = 3, and our exponential backoff has an initial delay of 100ms. It might be worth exploring in the future how we may want to more easily expose these configurations to users. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 9 个月前 | ||
| 8 个月前 | ||
| 8 个月前 | ||
| 8 个月前 | ||
| 9 个月前 | ||
| 1 年前 | ||
| 8 个月前 | ||
| 10 个月前 | ||
| 8 个月前 | ||
| 9 个月前 | ||
| 8 个月前 | ||
| 9 个月前 |