| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
chore: Drop Python 3.9 (#5479) ## Changes Made Python 3.9 is EOL starting from Nov 1, so lets drop it and make 3.10 the minimum. Also, start testing ranges 3.10 to 3.13 --------- Co-authored-by: Colin Ho <colin.ho99@gmail.com> | 10 个月前 | |
feat: Add a Bigtable data sink (#5431) ## Changes Made Allow dataframes to be written to Google Cloud Bigtable. E.g. >>> import daft >>> df = daft.from_pydict({"id": ["r5", "r6"], "my_col": ["hello", "world"]}) >>> for x in df.write_bigtable("desmond-project-475721", "desmond-instance", "desmond-table", "id", {"my_col": "cf1"}).collect(): ... print(x) ... {'write_responses': WriteResult(result={'status': 'success', 'rows_written': 2}, bytes_written=48, rows_written=2)} And if we later retrieve results in BigTable via: from google.cloud import bigtable client = bigtable.Client(project="my-project") instance = client.instance("my-instance") table = instance.table("my-table") for row in table.read_rows(): print(row.row_key) print(row.cells) print() we get b'r5' OrderedDict([('cf1', OrderedDict([(b'my_col', [<Cell value=b'hello' timestamp=2025-10-21 19:32:31.991000+00:00>])]))]) b'r6' OrderedDict([('cf1', OrderedDict([(b'my_col', [<Cell value=b'world' timestamp=2025-10-21 19:32:31.991000+00:00>])]))]) | 11 个月前 | |
docs: adds ai functions, ai providers, contributing, and docstrings with nav (#5438) ## Changes Made - Creates new AI section in Nav - Adds dedicated pages in User Guide for AI Functions and Providers usage - Moved how to add a new ai function page to new AI section - Adds examples to AI Function docstrings ## Related Issues Closes https://github.com/Eventual-Inc/Daft/issues/5399 ## Checklist - [x] Documented in API Docs (if applicable) - [x] Documented in User Guide (if applicable) - [x] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [x] Documentation builds and is formatted properly | 11 个月前 | |
chore: bump mypy and ruff in pre-commit (#5836) ## Changes Made Bump mypy to 1.19.1 and ruff to 0.14.10 and resolve pre-commit errors. Rationale: my local ruff version is much higher than that in the project pre-commit, so I am getting a lot of linter warnings in my IDE. It is also about time to upgrade lints to utilize newer Python language features. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
chore: bump mypy and ruff in pre-commit (#5836) ## Changes Made Bump mypy to 1.19.1 and ruff to 0.14.10 and resolve pre-commit errors. Rationale: my local ruff version is much higher than that in the project pre-commit, so I am getting a lot of linter warnings in my IDE. It is also about time to upgrade lints to utilize newer Python language features. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
fix: handle FileNotFoundError in read_huggingface fallback (#5831) ## Summary When HuggingFace's parquet API returns no files (glob returns no matches), read_huggingface now falls back to the datasets library instead of raising FileNotFoundError. This fixes the failing test_read_huggingface_datasets_doesnt_fail integration test caused by HuggingFace's parquet conversion being unavailable for some datasets. ## Root Cause The huggingface/documentation-images dataset's parquet API started returning empty results ({}) around 2025-12-17 14:30 UTC. This causes read_parquet to fail with FileNotFoundError when globbing for parquet files. The existing fallback only handled DaftCoreException with "Status(400" errors, not FileNotFoundError. ## CI Failure Details | Run ID | Time (UTC) | Status | Error | |--------|------------|--------|-------| | [20306131203](https://github.com/Eventual-Inc/Daft/actions/runs/20306131203) | 14:25:15 | ✅ success | (last pass) | | [20306481633](https://github.com/Eventual-Inc/Daft/actions/runs/20306481633) | 14:36:45 | ❌ failure | FileNotFoundError | | [20310015416](https://github.com/Eventual-Inc/Daft/actions/runs/20310015416) | 16:33:02 | ❌ failure | FileNotFoundError | | [20310356377](https://github.com/Eventual-Inc/Daft/actions/runs/20310356377) | 16:44:55 | ❌ failure | FileNotFoundError | Example error from CI: FileNotFoundError: File: hf://datasets/huggingface/documentation-images not found Glob path had no matches: "hf://datasets/huggingface/documentation-images". ## Changes - Added FileNotFoundError handling in read_huggingface to trigger fallback to datasets library - Refactored fallback logic into _fallback_to_datasets_library() helper to avoid duplication ## Test Plan - [x] Reproduced the failure locally before the fix - [x] Verified all 7 HuggingFace integration tests pass with the fix: - test_read_huggingface_datasets_doesnt_fail - test_read_huggingface[Eventual-Inc/sample-parquet-train-foo] - test_read_huggingface[fka/awesome-chatgpt-prompts-train-act] - test_read_huggingface[nebius/SWE-rebench-test-instance_id] - test_read_huggingface[SWE-Gym/SWE-Gym-train-instance_id] - test_read_huggingface_fallback_on_400_error - test_read_huggingface_multi_split_dataset | 9 个月前 | |
chore: bump mypy and ruff in pre-commit (#5836) ## Changes Made Bump mypy to 1.19.1 and ruff to 0.14.10 and resolve pre-commit errors. Rationale: my local ruff version is much higher than that in the project pre-commit, so I am getting a lot of linter warnings in my IDE. It is also about time to upgrade lints to utilize newer Python language features. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
fix: Optimize the small files issue of sink lance (#5844) ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> https://github.com/Eventual-Inc/Daft/issues/5843 Co-authored-by: cancai <caican@xiaomi.com> | 8 个月前 | |
feat(mcap): support topic_start_time_resolver and raw-bytes non-seekable reader (#5886) ## Changes Made * Add per-file per-topic start_time resolver and read via iter_messages with decoders disabled. * Use non-seekable stream to avoid seeking differences across FS backends. <!-- Describe what changes were made and why. Include implementation details if necessary. --> ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 8 个月前 | |
feat!: Catch transient errors on turbopuffer writes (#5380) ## Changes Made The turbopuffer client library [retries transient errors 4 times with appropriate back-offs and jitters](https://github.com/turbopuffer/turbopuffer-python?tab=readme-ov-file#retries). In the event that we receive a transient error from the client library, instead of causing the entire script to fail, we should simply track these transient errors and report the number of failed rows later on. Non-transient errors should cause a script to fail fast, as it did before. In the future, this should ideally extend to the proposal of splitting results so that we keep track of what rows resulted in errors so that we can retry them later. For now we simply track the number of failed rows. | 11 个月前 | |
feat: Add Apache Gravitino virtual file system (gvfs://) read support in io module (#5766) ## Changes Made Add Gravitino virtual file system (gvfs://) support read support. So that user can use "gvfs://filesets/catalog/schema/fileset_name" path to access s3 (and other cloud storages) location. This PR only implements the s3 as the first step. To ensure the functionality, the integration test cases are added with a MinIO service to simulate s3. ## Related Issues This is the second pr for https://github.com/Eventual-Inc/Daft/issues/5503 | 8 个月前 | |
chore: optimize operator naming (#5204) In the current progress, if the scan is of the Python version, its display is not user-friendly. Therefore, this PR attempts to fix this issue. The main approach is to convert PythonFunction into a specific Datasource Scan, for example Before PR tests/io/lancedb/test_lancedb_reads.py 🗡️ 🐟[1/1] ⠁ PythonFunction Scan | [00:00:00] 🗡️ 🐟[1/1] ⠁ PythonFunction Scan | [00:00:00] 2 rows out, 0 B bytes read 🗡️ 🐟[1/1] ✓ PythonFunction Scan | [00:00:00] 2 rows out, 0 B bytes read After PR tests/io/lancedb/test_lancedb_reads.py 🗡️ 🐟[1/1] ⠁ Lance(Python) Scan | [00:00:00] 🗡️ 🐟[1/1] ⠁ Lance(Python) Scan | [00:00:00] 2 rows out, 0 B bytes read 🗡️ 🐟[1/1] ✓ Lance(Python) Scan | [00:00:00] 2 rows out, 0 B bytes read ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly (tag @/ccmao1130 for docs review) | 11 个月前 | |
chore: bump mypy and ruff in pre-commit (#5836) ## Changes Made Bump mypy to 1.19.1 and ruff to 0.14.10 and resolve pre-commit errors. Rationale: my local ruff version is much higher than that in the project pre-commit, so I am getting a lot of linter warnings in my IDE. It is also about time to upgrade lints to utilize newer Python language features. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
chore: bump mypy and ruff in pre-commit (#5836) ## Changes Made Bump mypy to 1.19.1 and ruff to 0.14.10 and resolve pre-commit errors. Rationale: my local ruff version is much higher than that in the project pre-commit, so I am getting a lot of linter warnings in my IDE. It is also about time to upgrade lints to utilize newer Python language features. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
chore: bump mypy and ruff in pre-commit (#5836) ## Changes Made Bump mypy to 1.19.1 and ruff to 0.14.10 and resolve pre-commit errors. Rationale: my local ruff version is much higher than that in the project pre-commit, so I am getting a lot of linter warnings in my IDE. It is also about time to upgrade lints to utilize newer Python language features. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
chore: bump mypy and ruff in pre-commit (#5836) ## Changes Made Bump mypy to 1.19.1 and ruff to 0.14.10 and resolve pre-commit errors. Rationale: my local ruff version is much higher than that in the project pre-commit, so I am getting a lot of linter warnings in my IDE. It is also about time to upgrade lints to utilize newer Python language features. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
chore: Remove the old Ray Runner (#5375) ## Changes Made 🎉🎂🥳 Can finally delete it, Flotilla supports all necessary features. Additional Related Features Removed: * DataFrame.num_partitions: This is computed using the old Ray runner's planner. We could use the new Ray runner, but it seems kind of unnecessary * Context Settings: --------- Co-authored-by: Colin Ho <colin.ho99@gmail.com> | 10 个月前 | |
chore: bump mypy and ruff in pre-commit (#5836) ## Changes Made Bump mypy to 1.19.1 and ruff to 0.14.10 and resolve pre-commit errors. Rationale: my local ruff version is much higher than that in the project pre-commit, so I am getting a lot of linter warnings in my IDE. It is also about time to upgrade lints to utilize newer Python language features. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
chore: bump mypy and ruff in pre-commit (#5836) ## Changes Made Bump mypy to 1.19.1 and ruff to 0.14.10 and resolve pre-commit errors. Rationale: my local ruff version is much higher than that in the project pre-commit, so I am getting a lot of linter warnings in my IDE. It is also about time to upgrade lints to utilize newer Python language features. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
chore: Upgrade Ruff ruleset to 3.9 and add from __future__ import annotations (#4393) | 1 年前 | |
chore: bump mypy and ruff in pre-commit (#5836) ## Changes Made Bump mypy to 1.19.1 and ruff to 0.14.10 and resolve pre-commit errors. Rationale: my local ruff version is much higher than that in the project pre-commit, so I am getting a lot of linter warnings in my IDE. It is also about time to upgrade lints to utilize newer Python language features. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
feat: Add Common Crawl dataset (#5244) ## Changes Made Provide a simple, ergonomic way to access Common Crawl from Daft, so users can write: >>> import daft >>> daft.datasets.common_crawl("CC-MAIN-2025-33").show() ╭────────────────────────────────┬────────────────────────────────┬───────────┬─────────────────────────────────────────┬────────────────┬──────────────────────────────┬────────────────────────────────┬────────────────────────────────╮ │ WARC-Record-ID ┆ WARC-Target-URI ┆ WARC-Type ┆ WARC-Date ┆ Content-Length ┆ WARC-Identified-Payload-Type ┆ warc_content ┆ warc_headers │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ Utf8 ┆ Utf8 ┆ Utf8 ┆ Timestamp(Nanoseconds, Some("Etc/UTC")) ┆ Int64 ┆ Utf8 ┆ Binary ┆ Utf8 │ ╞════════════════════════════════╪════════════════════════════════╪═══════════╪═════════════════════════════════════════╪════════════════╪══════════════════════════════╪════════════════════════════════╪════════════════════════════════╡ │ 0313e9e8-9489-444e-bb20-ba477… ┆ None ┆ warcinfo ┆ 2025-09-05 11:21:01 UTC ┆ 492 ┆ None ┆ b"isPartOf: CC-MAIN-2025-38\r… ┆ {"Content-Type":"application/… │ ├╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┤ │ 53773031-a751-41c3-a7c5-b327a… ┆ http://0481.jp/g/tukuba/perfo… ┆ request ┆ 2025-09-05 13:07:28 UTC ┆ 320 ┆ None ┆ b"GET /g/tukuba/performance/1… ┆ {"Content-Type":"application/… │ ├╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┤ │ 4888f41c-0fb4-4ba4-afa0-a4a0a… ┆ http://0481.jp/g/tukuba/perfo… ┆ response ┆ 2025-09-05 13:07:28 UTC ┆ 16143 ┆ application/xhtml+xml ┆ b"HTTP/1.1 200 OK\r\nServer: … ┆ {"Content-Type":"application/… │ ├╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┤ │ 9eea2cef-8dfe-4d0d-8752-56d42… ┆ http://0481.jp/g/tukuba/perfo… ┆ metadata ┆ 2025-09-05 13:07:28 UTC ┆ 202 ┆ None ┆ b"fetchTimeMs: 783\r\ncharset… ┆ {"Content-Type":"application/… │ ├╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┤ │ 8c313289-61a9-4006-9924-824ce… ┆ http://0731yhwj.com/jjfa/1636… ┆ request ┆ 2025-09-05 12:16:36 UTC ┆ 327 ┆ None ┆ b"GET /jjfa/16365373911740375… ┆ {"Content-Type":"application/… │ ├╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┤ │ b8780f2e-eea4-4c8d-9e68-757bd… ┆ http://0731yhwj.com/jjfa/1636… ┆ response ┆ 2025-09-05 12:16:36 UTC ┆ 106231 ┆ text/html ┆ b"HTTP/1.1 200 OK\r\nDate: Fr… ┆ {"Content-Type":"application/… │ ├╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┤ │ 960529dc-840a-42df-b9f3-200c0… ┆ http://0731yhwj.com/jjfa/1636… ┆ metadata ┆ 2025-09-05 12:16:36 UTC ┆ 296 ┆ None ┆ b"fetchTimeMs: 696\r\ncharset… ┆ {"Content-Type":"application/… │ ├╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┤ │ 0657b53a-5691-41ed-bbf6-53d88… ┆ http://1-apple.com.tw/index.c… ┆ request ┆ 2025-09-05 11:49:26 UTC ┆ 401 ┆ None ┆ b"GET /index.cfm?Fuseaction=M… ┆ {"Content-Type":"application/… │ ╰────────────────────────────────┴────────────────────────────────┴───────────┴─────────────────────────────────────────┴────────────────┴──────────────────────────────┴────────────────────────────────┴────────────────────────────────╯ (Showing first 8 rows) >>> daft.datasets.common_crawl("CC-MAIN-2025-33", segment="1754151279521.11").limit(1).show() ╭────────────────────────────────┬─────────────────┬───────────┬─────────────────────────────────────────┬────────────────┬──────────────────────────────┬────────────────────────────────┬────────────────────────────────╮ │ WARC-Record-ID ┆ WARC-Target-URI ┆ WARC-Type ┆ WARC-Date ┆ Content-Length ┆ WARC-Identified-Payload-Type ┆ warc_content ┆ warc_headers │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ Utf8 ┆ Utf8 ┆ Utf8 ┆ Timestamp(Nanoseconds, Some("Etc/UTC")) ┆ Int64 ┆ Utf8 ┆ Binary ┆ Utf8 │ ╞════════════════════════════════╪═════════════════╪═══════════╪═════════════════════════════════════════╪════════════════╪══════════════════════════════╪════════════════════════════════╪════════════════════════════════╡ │ 526c37b2-f535-4015-b8dd-bfa8e… ┆ None ┆ warcinfo ┆ 2025-08-02 22:09:07 UTC ┆ 489 ┆ None ┆ b"isPartOf: CC-MAIN-2025-33\r… ┆ {"Content-Type":"application/… │ ╰────────────────────────────────┴─────────────────┴───────────┴─────────────────────────────────────────┴────────────────┴──────────────────────────────┴────────────────────────────────┴────────────────────────────────╯ (Showing first 1 of 1 rows) >>> daft.datasets.common_crawl("CC-MAIN-2025-33", segment="1754151279521.11", content="metadata").limit(1).show() ╭────────────────────────────────┬─────────────────┬───────────┬─────────────────────────────────────────┬────────────────┬──────────────────────────────┬────────────────────────────────┬────────────────────────────────╮ │ WARC-Record-ID ┆ WARC-Target-URI ┆ WARC-Type ┆ WARC-Date ┆ Content-Length ┆ WARC-Identified-Payload-Type ┆ warc_content ┆ warc_headers │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ Utf8 ┆ Utf8 ┆ Utf8 ┆ Timestamp(Nanoseconds, Some("Etc/UTC")) ┆ Int64 ┆ Utf8 ┆ Binary ┆ Utf8 │ ╞════════════════════════════════╪═════════════════╪═══════════╪═════════════════════════════════════════╪════════════════╪══════════════════════════════╪════════════════════════════════╪════════════════════════════════╡ │ cfae7e3e-02b7-4e13-b94e-8e9dd… ┆ None ┆ warcinfo ┆ 2025-08-16 01:03:20 UTC ┆ 277 ┆ None ┆ b"Software-Info: ia-web-commo… ┆ {"Content-Type":"application/… │ ╰────────────────────────────────┴─────────────────┴───────────┴─────────────────────────────────────────┴────────────────┴──────────────────────────────┴────────────────────────────────┴────────────────────────────────╯ (Showing first 1 of 1 rows) >>> daft.datasets.common_crawl("CC-MAIN-2025-33", segment="1754151279521.11", content="wet").limit(1).show() ╭────────────────────────────────┬─────────────────┬───────────┬─────────────────────────────────────────┬────────────────┬──────────────────────────────┬────────────────────────────────┬────────────────────────────────╮ │ WARC-Record-ID ┆ WARC-Target-URI ┆ WARC-Type ┆ WARC-Date ┆ Content-Length ┆ WARC-Identified-Payload-Type ┆ warc_content ┆ warc_headers │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ Utf8 ┆ Utf8 ┆ Utf8 ┆ Timestamp(Nanoseconds, Some("Etc/UTC")) ┆ Int64 ┆ Utf8 ┆ Binary ┆ Utf8 │ ╞════════════════════════════════╪═════════════════╪═══════════╪═════════════════════════════════════════╪════════════════╪══════════════════════════════╪════════════════════════════════╪════════════════════════════════╡ │ 0cb039e8-d357-485f-95dd-cdfdb… ┆ None ┆ warcinfo ┆ 2025-08-16 01:03:20 UTC ┆ 370 ┆ None ┆ b"Software-Info: ia-web-commo… ┆ {"Content-Type":"application/… │ ╰────────────────────────────────┴─────────────────┴───────────┴─────────────────────────────────────────┴────────────────┴──────────────────────────────┴────────────────────────────────┴────────────────────────────────╯ (Showing first 1 of 1 rows) >>> daft.datasets.common_crawl("CC-MAIN-2025-33", num_files=1).limit(1).show() ╭────────────────────────────────┬─────────────────┬───────────┬─────────────────────────────────────────┬────────────────┬──────────────────────────────┬────────────────────────────────┬────────────────────────────────╮ │ WARC-Record-ID ┆ WARC-Target-URI ┆ WARC-Type ┆ WARC-Date ┆ Content-Length ┆ WARC-Identified-Payload-Type ┆ warc_content ┆ warc_headers │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ Utf8 ┆ Utf8 ┆ Utf8 ┆ Timestamp(Nanoseconds, Some("Etc/UTC")) ┆ Int64 ┆ Utf8 ┆ Binary ┆ Utf8 │ ╞════════════════════════════════╪═════════════════╪═══════════╪═════════════════════════════════════════╪════════════════╪══════════════════════════════╪════════════════════════════════╪════════════════════════════════╡ │ 526c37b2-f535-4015-b8dd-bfa8e… ┆ None ┆ warcinfo ┆ 2025-08-02 22:09:07 UTC ┆ 489 ┆ None ┆ b"isPartOf: CC-MAIN-2025-33\r… ┆ {"Content-Type":"application/… │ ╰────────────────────────────────┴─────────────────┴───────────┴─────────────────────────────────────────┴────────────────┴──────────────────────────────┴────────────────────────────────┴────────────────────────────────╯ (Showing first 1 of 1 rows) See https://github.com/Eventual-Inc/Daft/discussions/5248 for more discussions and followups. | 1 年前 | |
chore: bump mypy and ruff in pre-commit (#5836) ## Changes Made Bump mypy to 1.19.1 and ruff to 0.14.10 and resolve pre-commit errors. Rationale: my local ruff version is much higher than that in the project pre-commit, so I am getting a lot of linter warnings in my IDE. It is also about time to upgrade lints to utilize newer Python language features. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
feat: Add Apache Gravitino virtual file system (gvfs://) write support in io module (#5965) | 8 个月前 | |
feat(lance): row-level schema evolution support (#5749) ## Changes Made ## Overview The following diagram describes the core workflow of the new add_columns feature. When a user calls write_lance(mode='merge'), the system checks for schema changes: - No new columns: Falls back to the existing append mode. - New columns present: Triggers the new row-level merge workflow, which groups data by fragment_id and efficiently merges new columns within each fragment using a join key (e.g., _rowaddr), ultimately committing only metadata changes. <img width="1016" height="2786" alt="whiteboard_exported_image-已去除(lightpdf cn)" src="https://github.com/user-attachments/assets/c7f2a3c6-ac3b-419e-bb61-7d5e540d302e" /> # Core Interface Changes This optimization introduces new APIs and extends existing interfaces. The core changes are as follows: ## 1. DataFrame.write_lance() The write_lance method adds a merge mode to implement efficient schema evolution. - Mode: mode: Literal["create", "append", "overwrite", "merge"] = "create" - New merge mode behavior: 1. Dataset does not exist: Behaves like create mode. 2. Dataset exists but no new columns: Behaves like append mode. 3. Dataset exists with new columns: Triggers the row-level merge workflow. In this mode, the DataFrame must contain fragment_id and a join key (defaults to _rowaddr or as specified by left_on/right_on). - UT Example (tests/io/lancedb/test_lance_merge_evolution.py): # Read existing data, including fragment_id and _rowaddr df_loaded = daft.read_lance( lance_dataset_path, default_scan_options={"with_row_address": True}, include_fragment_id=True ) # Derive a new column df_with_new_col = df_loaded.with_column("double_lat", daft.col("lat") * 2) # Write with merge mode; Daft handles column merging automatically df_with_new_col.write_lance(lance_dataset_path, mode="merge") ## 2. daft.io.lance.read_lance() read_lance adds the include_fragment_id parameter to expose each fragment's ID as a new column (fragment_id) to the user upon reading. This is a key prerequisite for row-level merging. - New Parameter: include_fragment_id: Optional[bool] = None - UT Example (tests/io/lancedb/test_lance_merge_evolution.py): df = daft.read_lance( lance_dataset_path, default_scan_options={"with_row_address": True}, include_fragment_id=True ) assert "fragment_id" in df.column_names ## 3. daft.io.lance.merge_columns_df() This is a new low-level API that provides a more flexible, DataFrame-based row-level column merge capability. write_lance(mode='merge') internally wraps this function. - Core Functionality: Receives a DataFrame containing fragment_id, a join key, and new columns, and merges it efficiently into an existing Lance dataset. - UT Example (tests/io/lancedb/test_lance_merge_evolution.py): # Prepare a DataFrame with only fragment_id, join key, and new columns df_subset = df_loaded.with_column("double_lat", daft.col("lat") * 2)\ .select("_rowaddr", "fragment_id", "double_lat") # Call the low-level API to perform the merge daft.io.lance.merge_columns_df( df_subset, lance_dataset_path, read_columns=["_rowaddr", "double_lat"], ) ## 4. Interface Comparison Both interfaces support column merging, but their purpose and usage differ: - Interface 1: DataFrame.write_lance(mode="merge") — A writer-side merge; automatically detects new columns; requires fragment_id and a join key; returns write statistics metadata. - Interface 2: daft.io.lance.merge_columns_df() — A DataFrame-based row-level merge; explicitly takes a DataFrame with only fragment_id, join key, and new columns; returns None; allows for finer control via read_columns and batch_size. Usage Recommendation: - For end-to-end write pipelines with more intuitive semantics, prefer write_lance(mode="merge"). - For complex DataFrame transformations or custom column reading before writing, use merge_columns_df(). # Deprecation Plan With the introduction of the new write_lance(mode='merge') and daft.io.lance.merge_columns_df, the old daft.io.lance.merge_columns interface has been marked as deprecated. - Reason for Deprecation: The old interface was based on a fragment-level transformation function, which did not support DataFrames that had undergone complex operations (like joins, groupbys, etc.), limiting its flexibility and use cases. - Future Plan: daft.io.lance.merge_columns will be removed in a future release. All users are advised to migrate to the new write_lance(mode='merge') interface, which provides a more powerful and user-friendly column extension capability. <!-- Describe what changes were made and why. Include implementation details if necessary. --> ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 9 个月前 | |
fix: Enable strict mode for mypy in pre-commit (#4422) ## Changes Made Fixes all the remaining mypy errors and enables strict mode in pre-commit. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly (tag @/ccmao1130 for docs review) | 1 年前 | |
feat: adds column set visitor and use in pushdowns (#4929) ## Changes Made * Adds a visitor which collects a column set * Uses this in Pushdowns to return the filter column set. ## Related Issues I considered saving the list from the rust pushdowns object, but a python visitor is more generic and flexible because it can be used outside the context of pushdowns to find all the columns in an arbitrary expression. This also uses a set for unique column names. If/when daft has a notion of bound columns, then the hash will be important for distinctness because for now we rely on unique names within an expression tree. ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly (tag @/ccmao1130 for docs review) | 1 年前 | |
feat(optimizer): Add Lance count() pushdown optimization (#4969) Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> | 1 年前 | |
fix: Reraise exceptions from data sinks to handle unserializable exceptions (#4858) ## Changes Made If an exception is thrown in the data sink and is unserializable, we may get errors like File ~/anaconda3/lib/python3.12/site-packages/daft/dataframe/dataframe.py:1304, in DataFrame.write_sink(self, sink) 1302 builder = self._builder.write_datasink(sink.name(), sink) 1303 write_df = DataFrame(builder) -> 1304 write_df.collect() 1306 results = write_df.to_pydict() 1307 assert "write_results" in results ... ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/ray/anaconda3/lib/python3.12/site-packages/ray/exceptions.py", line 54, in from_ray_exception raise RuntimeError(msg) from e RuntimeError: Failed to unpickle serialized exception Catching and reraising their string representations allow us to report the actual exception to the user. | 1 年前 | |
feat: adds python partition fields to the DataSource API (#4449) ## Changes Made Now that we have python classes for PartitionField and PartitionSpec, they can be hooked up to the DataSource API, and the __shim does the PyPartitionField unwrapping. ## Related Issues - #4228 ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly (tag @/ccmao1130 for docs review) | 1 年前 | |
| 8 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 10 个月前 | ||
| 11 个月前 | ||
| 11 个月前 | ||
| 9 个月前 | ||
| 9 个月前 | ||
| 9 个月前 | ||
| 9 个月前 | ||
| 8 个月前 | ||
| 8 个月前 | ||
| 11 个月前 | ||
| 8 个月前 | ||
| 11 个月前 | ||
| 9 个月前 | ||
| 9 个月前 | ||
| 9 个月前 | ||
| 9 个月前 | ||
| 10 个月前 | ||
| 9 个月前 | ||
| 9 个月前 | ||
| 1 年前 | ||
| 9 个月前 | ||
| 1 年前 | ||
| 9 个月前 | ||
| 8 个月前 | ||
| 9 个月前 | ||
| 1 年前 | ||
| 1 年前 | ||
| 1 年前 | ||
| 1 年前 | ||
| 1 年前 | ||
| 8 个月前 |