| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
[Ray Datasets] Port Ray Datasets integration. (#837) This PR ports the Ray Datasets integration to the new backend, including representing the tensor extension type with our Python object type and support Datasets with block types other than Arrow. ## Drivebys - Also adds a DataFrame.to_arrow() API, as an analog for the existing DataFrame.to_pandas() API. Having this available made writing the roundtrip tests much easier. | 3 年前 | |
feat: support label_selector specification in ray actor/task creation (#5042) ## Changes Made When there are multiple worker groups in a Ray cluster, tasks need to be submitted based on the specific labels of the worker groups. Therefore, a label_selector parameter has been added to the UDF. For detailed usage, please refer to the Ray documentation at https://docs.ray.io/en/latest/ray-core/scheduling/labels.html#monitor-nodes-using-labels <!-- Describe what changes were made and why. Include implementation details if necessary. --> ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly (tag @/ccmao1130 for docs review) | 9 个月前 | |
chore: Remove the old Ray Runner (#5375) ## Changes Made 🎉🎂🥳 Can finally delete it, Flotilla supports all necessary features. Additional Related Features Removed: * DataFrame.num_partitions: This is computed using the old Ray runner's planner. We could use the new Ray runner, but it seems kind of unnecessary * Context Settings: --------- Co-authored-by: Colin Ho <colin.ho99@gmail.com> | 10 个月前 | |
chore(deps): drop pyarrow 8.0.0 support, bump minimum to >= 15.0.0 (#6378) ## Changes Made Bump minimum PyArrow version from >= 8.0.0 to >= 15.0.0 (44 files, -450 lines). - Update version constraints in pyproject.toml and CI test matrix - Remove _FixSliceOffsets workaround (pyarrow < 12.0.0 struct slice offset bug, fixed upstream) - Remove pyarrow_supports_fixed_shape_tensor() and all conditional branches - Remove obsolete version checks (< 12.0.1, < 13.0.0, >= 9.0.0) and try/except imports - Clean up ~50 pytest.mark.skipif markers and unused PYARROW_GE_* constants in test files Note: _FixEmptyStructArrays is intentionally kept — Daft internally cannot handle empty StructArrays, not just an arrow2 FFI issue. ## Related Issues Closes #6347 | 6 个月前 | |
feat: support to_ray_dataset() from the native runner (#6486) DataFrame.to_ray_dataset() previously required the Ray runner, raising a ValueError when called from the native runner. This forced users who run Daft compute on the native runner (for performance) but need Ray Datasets for downstream ML pipelines to do manual Arrow/pylist round-trips. The restriction was artificial. When using the native runner, partitions can be converted to Arrow tables and passed directly to Ray viaray.data.from_arrow(). --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> | 6 个月前 | |
fix(ray): namespace flotilla actor per job to avoid plan id collisions (#5855) Namespace the RemoteFlotillaRunner Ray actor name per Ray job / process to avoid cross-client plan ID collisions when multiple Daft clients connect to the same Ray cluster. - Introduce a cached suffix used to namespace the flotilla plan runner actor name per Ray job / process. - Derive the suffix from the Ray runtime context job_id when available. - If no job identifier can be determined, fall back to a per-process UUID so the actor name remains unique across different Python processes. - Update FlotillaRunner to construct the Ray actor with the new per-job actor name while keeping the namespace="daft" and get_if_exists=True semantics unchanged. - Add comments explaining the root cause: DistributedPhysicalPlan / QueryIdx uses a per-process counter, so plan IDs like "0" can collide across different clients if they share the same named actor instance. - Add a lightweight unit test that verifies the flotilla actor name is derived from the base name plus a job-specific suffix and remains stable within a process. ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> ## Related Issues #5856 <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
chore(tests): migrate internal usages of daft.udf to cls/func (#6348) update tests to remove a bunch of deprecation warnings in CI/tests ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 6 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 3 年前 | ||
| 9 个月前 | ||
| 10 个月前 | ||
| 6 个月前 | ||
| 6 个月前 | ||
| 9 个月前 | ||
| 6 个月前 |