| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
chore: Remove the old Ray Runner (#5375) ## Changes Made 🎉🎂🥳 Can finally delete it, Flotilla supports all necessary features. Additional Related Features Removed: * DataFrame.num_partitions: This is computed using the old Ray runner's planner. We could use the new Ray runner, but it seems kind of unnecessary * Context Settings: --------- Co-authored-by: Colin Ho <colin.ho99@gmail.com> | 9 个月前 | |
fix: overriding dimensions for openai embedding models (#6013) ## Changes Made When overriding the embedding dimension to embed_text with the OpenAI provider, it would fail with this message: daft.exceptions.DaftTypeError: Can not cast Tensor array to FixedShapeTensor array with type FixedShapeTensor(Float32, [256]): Tensor array has shapes different than [256]; This was because we were only passing in the dimensions when embed_options["supports_overriding_dimensions"] was true, but that was not set, even though we hardcoded the support in the _models variable. This PR sets the supports_overriding_dimensions option when the OpenAI model supports it, fixing the issue. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 8 个月前 | |
feat: audio file subtype (#5602) ## Changes Made Creates new AudioFile class that's mostly a wrapper around soundfile library. provides common methods for interacting with audio files such as - metadata - resample - to_numpy Also supports dataframe expressions - daft.functions.audio_file - daft.functions.resample - daft.functions.audio_metadata Still need to add tests and a few more methods. Will follow up in a separate PR some functionality for audio extraction from video files. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly | 10 个月前 | |
chore: Remove runner from context (#5628) ## Changes Made Remove runner methods from daft.context, as they are now decoupled from the context. Changed all uses of daft.context.set_runner_... to daft.set_runner_... ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly | 9 个月前 | |
feat: Support pattern filtering for SHOW TABLES (#5423) ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> ### Extended SQL LIKE Pattern Matching in MemoryCatalog - Upgrade sqlparser to extract namespace from query - Translate SQL LIKE patterns to regex (src/daft-catalog/src/pattern.rs) - Supports standard SQL LIKE wildcards: % (zero or more), _ (exactly one), \ (escape) - Updated documentation on SHOW TABLES syntax and pattern behavior ### Testing - Added pattern matching tests to test_sql_show_tables.py and in src/daft-catalog/pattern.rs - cargo test -p daft-catalog pattern:: ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> Closes #4461 Closes #4007 ## Checklist - [x] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly | 9 个月前 | |
fix: Set default ImageMode in decode_image to RGB (#5827) ## Changes Made Set the default ImageMode for decode_image to RGB. Currently, the behavior is to infer the image mode at a per-image level, which means we could have 8-bit and 16-bit images in the same column, which will error because the underlying datatypes (uint8 vs uint16) are different. The solution to this is to force decode into a specific image mode. This PR simply makes that the default. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
fix(observability): Clean up progress bar naming (#6028) ## Changes Made While working on the One True Progress Bar (™️ pending) PR, noticed a couple of things we can clean up. | 8 个月前 | |
feat: Explicit AWS vs. HTTP mode for common crawl dataset (#5379) Adds a new required argument to daft.datasets.common_crawl: in_aws: bool. This **must** be set to True when running in AWS and False when running outside of AWS. This allows Daft to select the most optimal download strategy for CC data. Added a notice about this to the docstring. Refactors the existing mocked unit tests for this by making the tests patch the appropriate _get_{s3,http}_manifest_path using the value of in_aws. Adds in_aws as a pytest parameter and parameterizes each test on True and False. Updates the Common Crawl documentation to mention the new required in_aws parameter. Adds a new section discussing the new HTTP download mode and provides an example. | 11 个月前 | |
feat: Add guess_mime_type scalar expression for MIME type detection from bytes (#5883) | 8 个月前 | |
fix(video): correct keyframe seek timestamp calculation for start_time (#6005) ## Changes Made VideoFile.keyframes() computed an incorrect seek timestamp when start_time > 0. The previous implementation used start_time * time_base, which keeps the effective seek position near 0 for small time_base values (for example 1/30000). As a result, the start_time parameter was effectively ignored and keyframes were decoded from the beginning of the video. This also affected the daft.functions.video_keyframes() UDF, which delegates to VideoFile.keyframes(). Users expecting to extract keyframes from a later time range (for example starting at 10 seconds) instead received frames starting close to time 0. <!-- Describe what changes were made and why. Include implementation details if necessary. --> ## Related Issues #5949 <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 8 个月前 | |
chore: Remove expression namespaces (#5619) ## Changes Made Bye bye ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly | 9 个月前 | |
feat: Add Apache Gravitino virtual file system (gvfs://) write support in io module (#5965) | 8 个月前 | |
feat(mcap): support topic_start_time_resolver and raw-bytes non-seekable reader (#5886) ## Changes Made * Add per-file per-topic start_time resolver and read via iter_messages with decoders disabled. * Use non-seekable stream to avoid seeking differences across FS backends. <!-- Describe what changes were made and why. Include implementation details if necessary. --> ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 8 个月前 | |
fix: Optimize the display information of Join nodes in query plan (#5617) ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> Currently, when displaying the query plan containing a "join" node using explain, we can't intuitively distinguish between the left and right tables, especially in the case of a Broadcast Join, which requires the is_swapped field to assist in judgment. This PR mainly unifies the information format and elements of the Join nodes displayed in explain, and explicitly identifies the Broadcaster and Receiver for Broadcast Join. For example: text == Physical Plan == * BroadcastJoin | Type: Inner | Left: Join key = col(1: l_name), Role = Receiver | Right: Join key = col(1: s_name), Role = Broadcaster | Null equals nulls: [false] |\ | * ScanTaskSource: | | Num Scan Tasks = 100 | | Estimated Scan Bytes = 728327 | | Pushdowns: {filter: not(is_null(col(l_name)))} | | Schema: {id#Int64, l_name#Utf8, l_email#Utf8} | | Scan Tasks: [ ... ] | * Project: col(0: id) as right.id, col(1: s_name), col(2: s_email) | Resource request = None | * ScanTaskSource: | Num Scan Tasks = 10 | Estimated Scan Bytes = 67702 | Pushdowns: {filter: not(is_null(col(s_name)))} | Schema: {id#Int64, s_name#Utf8, s_email#Utf8} | Scan Tasks: [ ... ] ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly Signed-off-by: plotor <zhenchao.wang@hotmail.com> | 8 个月前 | |
chore: Upgrade Ruff ruleset to 3.9 and add from __future__ import annotations (#4393) | 1 年前 | |
chore: bump mypy and ruff in pre-commit (#5836) ## Changes Made Bump mypy to 1.19.1 and ruff to 0.14.10 and resolve pre-commit errors. Rationale: my local ruff version is much higher than that in the project pre-commit, so I am getting a lot of linter warnings in my IDE. It is also about time to upgrade lints to utilize newer Python language features. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 8 个月前 | |
feat: Support configuring Conda Env for Class UDFs in Flotilla (#5117) ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> #5104 ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly (tag @/ccmao1130 for docs review) Signed-off-by: plotor <zhenchao.wang@hotmail.com> | 8 个月前 | |
| 8 个月前 | ||
chore: bump mypy and ruff in pre-commit (#5836) ## Changes Made Bump mypy to 1.19.1 and ruff to 0.14.10 and resolve pre-commit errors. Rationale: my local ruff version is much higher than that in the project pre-commit, so I am getting a lot of linter warnings in my IDE. It is also about time to upgrade lints to utilize newer Python language features. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 8 个月前 | |
fix: ".*" not handled correctly in SQL planner (#5784) ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> 1. Fix handling of struct field wildcards (.*) in the SQL planner. 2. Add test case to validate. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> https://github.com/Eventual-Inc/Daft/issues/4120 | 8 个月前 | |
fix: Supporting fractional gpu count on class udf (#5840) ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> Supporting fractional gpu count in class udf. def cls( class_: type | None = None, *, gpus: float = 0.0, use_process: bool | None = None, max_concurrency: int | None = None, max_retries: int | None = None, on_error: Literal["raise", "log", "ignore"] | None = None, ) -> type | Callable[[type], type]: before this, the datatype of gpus in @daft.cls was int, but the underlying rust implementation was float. The declarations on the python side and the rust side are inconsistent. And for small models, we should support fine-grained resource Settings, such as 0.5 gpu. Declarations of the float type make it clearer and accurate. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> https://github.com/Eventual-Inc/Daft/issues/5837 --------- Co-authored-by: cancai <caican@xiaomi.com> Co-authored-by: Colin Ho <chiuhong@usc.edu> | 8 个月前 | |
chore: Remove expression namespaces (#5619) ## Changes Made Bye bye ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly | 9 个月前 | |
chore: optimizes slow tests in CI/CD (#6029) ## Results The PR reduced unit-test time (3.13, native, 22.0.0, ubuntu-latest, false) from [33m 18s](https://github.com/Eventual-Inc/Daft/actions/runs/20981828603/job/60308365885?pr=6023) to 13m 44s — saving about 20 mins which is nice. The test_context is still slow, but it has very few tests, a big win is parallelization and taking >5 mins off just by skipping the setup for skipped tests. This also reduced the ray tests from 45 mins to 27 mins, and there should be plenty more savings here. Then disabling windows eliminated a 50+ min action. text ================================================================================ OVERALL SUMMARY ================================================================================ Tests Analyzed: 9620 Total Time: 941.32s Average Time: 0.05s Max Time: 14.70s Min Time: 0.01s It took 188s to run make test on my mac which is nice; it also crushed my machine while it was running 🥲 ## Changes Made * Adds -n auto to pytest for parallelize testing on multiple cores * Skips generating fixtures (expensive?) by not actually creating the tests - we were creating test data _then_ skipping for several thousand join tests taking ~5 mins of setup work just to skip. * Memoizes data generation for the window tests (often the slowest for me) * Reduces the data generation size and faster comparison algo (before was 10,000 data points with a O(n^2) algo, now it's 1,000 data points with an O(nlogn) comparison algo) * Disables the RAY_DASHBOARD for the context setup commands, and other minor things to speedup ray init * Skips windows on PRs (>50 minutes) only enabled on main, worthwhile trade-off because we will still catch windows-specific failures when merging to main which is quite rare to be fair. ## Related Issues N/A | 8 个月前 | |
add tests | 4 年前 | |
chore: optimizes slow tests in CI/CD (#6029) ## Results The PR reduced unit-test time (3.13, native, 22.0.0, ubuntu-latest, false) from [33m 18s](https://github.com/Eventual-Inc/Daft/actions/runs/20981828603/job/60308365885?pr=6023) to 13m 44s — saving about 20 mins which is nice. The test_context is still slow, but it has very few tests, a big win is parallelization and taking >5 mins off just by skipping the setup for skipped tests. This also reduced the ray tests from 45 mins to 27 mins, and there should be plenty more savings here. Then disabling windows eliminated a 50+ min action. text ================================================================================ OVERALL SUMMARY ================================================================================ Tests Analyzed: 9620 Total Time: 941.32s Average Time: 0.05s Max Time: 14.70s Min Time: 0.01s It took 188s to run make test on my mac which is nice; it also crushed my machine while it was running 🥲 ## Changes Made * Adds -n auto to pytest for parallelize testing on multiple cores * Skips generating fixtures (expensive?) by not actually creating the tests - we were creating test data _then_ skipping for several thousand join tests taking ~5 mins of setup work just to skip. * Memoizes data generation for the window tests (often the slowest for me) * Reduces the data generation size and faster comparison algo (before was 10,000 data points with a O(n^2) algo, now it's 1,000 data points with an O(nlogn) comparison algo) * Disables the RAY_DASHBOARD for the context setup commands, and other minor things to speedup ray init * Skips windows on PRs (>50 minutes) only enabled on main, worthwhile trade-off because we will still catch windows-specific failures when merging to main which is quite rare to be fair. ## Related Issues N/A | 8 个月前 | |
chore: optimizes slow tests in CI/CD (#6029) ## Results The PR reduced unit-test time (3.13, native, 22.0.0, ubuntu-latest, false) from [33m 18s](https://github.com/Eventual-Inc/Daft/actions/runs/20981828603/job/60308365885?pr=6023) to 13m 44s — saving about 20 mins which is nice. The test_context is still slow, but it has very few tests, a big win is parallelization and taking >5 mins off just by skipping the setup for skipped tests. This also reduced the ray tests from 45 mins to 27 mins, and there should be plenty more savings here. Then disabling windows eliminated a 50+ min action. text ================================================================================ OVERALL SUMMARY ================================================================================ Tests Analyzed: 9620 Total Time: 941.32s Average Time: 0.05s Max Time: 14.70s Min Time: 0.01s It took 188s to run make test on my mac which is nice; it also crushed my machine while it was running 🥲 ## Changes Made * Adds -n auto to pytest for parallelize testing on multiple cores * Skips generating fixtures (expensive?) by not actually creating the tests - we were creating test data _then_ skipping for several thousand join tests taking ~5 mins of setup work just to skip. * Memoizes data generation for the window tests (often the slowest for me) * Reduces the data generation size and faster comparison algo (before was 10,000 data points with a O(n^2) algo, now it's 1,000 data points with an O(nlogn) comparison algo) * Disables the RAY_DASHBOARD for the context setup commands, and other minor things to speedup ray init * Skips windows on PRs (>50 minutes) only enabled on main, worthwhile trade-off because we will still catch windows-specific failures when merging to main which is quite rare to be fair. ## Related Issues N/A | 8 个月前 | |
chore: Pin dependencies (#5667) ## Changes Made Add upper bound to all core and optional dependencies to prevent breaking changes from affecting users. Additionally add == pins to all dev dependencies, for reproducible builds and tests. Upgrading dependencies can be done via dependabot. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
chore: Drop Python 3.9 (#5479) ## Changes Made Python 3.9 is EOL starting from Nov 1, so lets drop it and make 3.10 the minimum. Also, start testing ranges 3.10 to 3.13 --------- Co-authored-by: Colin Ho <colin.ho99@gmail.com> | 10 个月前 | |
feat: Add pow expression (#5237) ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> ## Related Issues #4704 <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly (tag @/ccmao1130 for docs review) --------- Co-authored-by: Colin Ho <colin.ho99@gmail.com> | 10 个月前 | |
feat: Better errors when lazy imports fail (#5753) Accessing an attribute on a lazy-imported module will now raise an ImportError with an informative error message, rather than a fairly cryptic AttributeError like the following. | 9 个月前 | |
fix: make it easier to enable different logging levels (#5661) ## Changes Made make it easier to enable different logging levels. Here's a summary of the PR. - Added convenient logger setup functions: setup_debug_logger, setup_info_logger, setup_warn_logger, and setup_error_logger for quickly configuring different log levels. - Centralized log level configuration logic in setup_logger_level, which supports setting the log level, filtering by prefix, and restricting logs to Daft module only. - When setting the log level, all handler filters are cleared before new filters are added, allowing flexible log filtering. - Added validation for log level input to ensure only valid levels are accepted. - After each logger setup, refresh_logger() is called to synchronize the configuration with the Rust backend. ## Related Issues Closes #5651 | 9 个月前 | |
ci: Remove Tests for the Old Ray Runner (#5374) Co-authored-by: Colin Ho <colin.ho99@gmail.com> | 11 个月前 | |
chore: bump mypy and ruff in pre-commit (#5836) ## Changes Made Bump mypy to 1.19.1 and ruff to 0.14.10 and resolve pre-commit errors. Rationale: my local ruff version is much higher than that in the project pre-commit, so I am getting a lot of linter warnings in my IDE. It is also about time to upgrade lints to utilize newer Python language features. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 8 个月前 | |
refactor: Cleanup Dtype Names (#5400) | 11 个月前 | |
feat: Support pattern filtering for SHOW TABLES (#5423) ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> ### Extended SQL LIKE Pattern Matching in MemoryCatalog - Upgrade sqlparser to extract namespace from query - Translate SQL LIKE patterns to regex (src/daft-catalog/src/pattern.rs) - Supports standard SQL LIKE wildcards: % (zero or more), _ (exactly one), \ (escape) - Updated documentation on SHOW TABLES syntax and pattern behavior ### Testing - Added pattern matching tests to test_sql_show_tables.py and in src/daft-catalog/pattern.rs - cargo test -p daft-catalog pattern:: ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> Closes #4461 Closes #4007 ## Checklist - [x] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly | 9 个月前 | |
chore: bump mypy and ruff in pre-commit (#5836) ## Changes Made Bump mypy to 1.19.1 and ruff to 0.14.10 and resolve pre-commit errors. Rationale: my local ruff version is much higher than that in the project pre-commit, so I am getting a lot of linter warnings in my IDE. It is also about time to upgrade lints to utilize newer Python language features. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 8 个月前 | |
feat: Use JSON Serialization for Plans in Subscribers (#5709) ## Changes Made Use our JSON plan serialization instead of the NodeInfo struct or our ASCII repr for sending plans to subscribers, including the dashboard. This will make it easier to support subscribers in Flotilla, since we don't need to use Swordfish-specific properties like whats in NodeInfo. Plus, the dashboard can now use / visualize additional fields in specific operators, like Project expressions. Tagging @samstokes and @ohbh for visibility on consuming the optimized plan. Feel free to review too | 9 个月前 | |
fix: Optimize the display information of Join nodes in query plan (#5617) ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> Currently, when displaying the query plan containing a "join" node using explain, we can't intuitively distinguish between the left and right tables, especially in the case of a Broadcast Join, which requires the is_swapped field to assist in judgment. This PR mainly unifies the information format and elements of the Join nodes displayed in explain, and explicitly identifies the Broadcaster and Receiver for Broadcast Join. For example: text == Physical Plan == * BroadcastJoin | Type: Inner | Left: Join key = col(1: l_name), Role = Receiver | Right: Join key = col(1: s_name), Role = Broadcaster | Null equals nulls: [false] |\ | * ScanTaskSource: | | Num Scan Tasks = 100 | | Estimated Scan Bytes = 728327 | | Pushdowns: {filter: not(is_null(col(l_name)))} | | Schema: {id#Int64, l_name#Utf8, l_email#Utf8} | | Scan Tasks: [ ... ] | * Project: col(0: id) as right.id, col(1: s_name), col(2: s_email) | Resource request = None | * ScanTaskSource: | Num Scan Tasks = 10 | Estimated Scan Bytes = 67702 | Pushdowns: {filter: not(is_null(col(s_name)))} | Schema: {id#Int64, s_name#Utf8, s_email#Utf8} | Scan Tasks: [ ... ] ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly Signed-off-by: plotor <zhenchao.wang@hotmail.com> | 8 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 9 个月前 | ||
| 8 个月前 | ||
| 10 个月前 | ||
| 9 个月前 | ||
| 9 个月前 | ||
| 9 个月前 | ||
| 8 个月前 | ||
| 11 个月前 | ||
| 8 个月前 | ||
| 8 个月前 | ||
| 9 个月前 | ||
| 8 个月前 | ||
| 8 个月前 | ||
| 8 个月前 | ||
| 1 年前 | ||
| 8 个月前 | ||
| 8 个月前 | ||
| 8 个月前 | ||
| 8 个月前 | ||
| 8 个月前 | ||
| 8 个月前 | ||
| 9 个月前 | ||
| 8 个月前 | ||
| 4 年前 | ||
| 8 个月前 | ||
| 8 个月前 | ||
| 9 个月前 | ||
| 10 个月前 | ||
| 10 个月前 | ||
| 9 个月前 | ||
| 9 个月前 | ||
| 11 个月前 | ||
| 8 个月前 | ||
| 11 个月前 | ||
| 9 个月前 | ||
| 8 个月前 | ||
| 9 个月前 | ||
| 8 个月前 |