| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
feat: add support for reading json arrays (#4844) ## Changes Made General approach is that we peek at the first byte of the file. if it's [ we use the array algorithm, if its { we use the default ndjson algo. ⚠️ It's worth noting that this is not highly optimized and _could_ OOM on very large JSON files as they are pulled **entirely** into memory. as a side note: Our JSON reading logic is a total mess and I didn't want to do any unnecessary refactors, so there was a bit of copy/paste in this PR to get things to work. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly (tag @/ccmao1130 for docs review) | 1 年前 | |
fix: Resolve mismatch with Thrift compact protocol (#4545) ## Changes Made The [Thrift compact protocol](https://github.com/apache/thrift/blob/master/doc/specs/thrift-compact-protocol.md) is used for Parquet file metadata. [parquet-format-safe](https://github.com/jorgecarleitao/parquet-format-safe) and other Rust implementations of the protocol eagerly read string/binary fields as UTF-8. However, based on the protocol which states that > Strings are first encoded to UTF-8, and then send as binary it cannot be known upfront, without using the schema to disambiguate the field type, whether a field is a string or a binary. This means that when the field is actually a binary field and contains invalid UTF-8, Rust libraries error out when reading the field with File out of specification: Invalid thrift: bad data. To fix this, we patch the protocol implementation to correctly interpret string/binary fields as binary. ## Related Issues Closes #4515 | 1 年前 | |
[FEAT] Add string tokenize expression (#2503) Allows users to tokenize a string column using tiktoken and a variety of encoders. Todo list: - [x] Support for builtin models (cl100k_base, p50k_base, etc) - [x] Support for loading models from a token file - [x] Support for downloading models from the cloud - [x] More tests - [x] Fix error handling - [x] Pattern argument - [x] Special token support - [x] All the tests - [x] Update docs Things that could be done in the future: - Add caching for token files so that the processes don't have to each download it. - Make a fork of tiktoken-rs for various fixes (accessing private fields, fixing unwraps, optimizing etc) - Add support for huggingface tokenizers - Add more granular special token support (custom inputs, using a subset of them) | 2 年前 | |
[CHORE] Add TPC-H questions 11-22 to benchmarks (currently skipped) (#2299) | 2 年前 | |
Initializes working rust-main branch * Old tests are migrated to tests-legacy/ - only new, working tests are left in tests/ * Tests that are failing in tests/ are skipped and tagged with [RUST-INT] * Small fixes made for the code to type-check - removal of udf.py and stubbing out of DataFrame.explode() code | 3 年前 | |
feat: Add Common Crawl dataset (#5244) ## Changes Made Provide a simple, ergonomic way to access Common Crawl from Daft, so users can write: >>> import daft >>> daft.datasets.common_crawl("CC-MAIN-2025-33").show() ╭────────────────────────────────┬────────────────────────────────┬───────────┬─────────────────────────────────────────┬────────────────┬──────────────────────────────┬────────────────────────────────┬────────────────────────────────╮ │ WARC-Record-ID ┆ WARC-Target-URI ┆ WARC-Type ┆ WARC-Date ┆ Content-Length ┆ WARC-Identified-Payload-Type ┆ warc_content ┆ warc_headers │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ Utf8 ┆ Utf8 ┆ Utf8 ┆ Timestamp(Nanoseconds, Some("Etc/UTC")) ┆ Int64 ┆ Utf8 ┆ Binary ┆ Utf8 │ ╞════════════════════════════════╪════════════════════════════════╪═══════════╪═════════════════════════════════════════╪════════════════╪══════════════════════════════╪════════════════════════════════╪════════════════════════════════╡ │ 0313e9e8-9489-444e-bb20-ba477… ┆ None ┆ warcinfo ┆ 2025-09-05 11:21:01 UTC ┆ 492 ┆ None ┆ b"isPartOf: CC-MAIN-2025-38\r… ┆ {"Content-Type":"application/… │ ├╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┤ │ 53773031-a751-41c3-a7c5-b327a… ┆ http://0481.jp/g/tukuba/perfo… ┆ request ┆ 2025-09-05 13:07:28 UTC ┆ 320 ┆ None ┆ b"GET /g/tukuba/performance/1… ┆ {"Content-Type":"application/… │ ├╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┤ │ 4888f41c-0fb4-4ba4-afa0-a4a0a… ┆ http://0481.jp/g/tukuba/perfo… ┆ response ┆ 2025-09-05 13:07:28 UTC ┆ 16143 ┆ application/xhtml+xml ┆ b"HTTP/1.1 200 OK\r\nServer: … ┆ {"Content-Type":"application/… │ ├╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┤ │ 9eea2cef-8dfe-4d0d-8752-56d42… ┆ http://0481.jp/g/tukuba/perfo… ┆ metadata ┆ 2025-09-05 13:07:28 UTC ┆ 202 ┆ None ┆ b"fetchTimeMs: 783\r\ncharset… ┆ {"Content-Type":"application/… │ ├╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┤ │ 8c313289-61a9-4006-9924-824ce… ┆ http://0731yhwj.com/jjfa/1636… ┆ request ┆ 2025-09-05 12:16:36 UTC ┆ 327 ┆ None ┆ b"GET /jjfa/16365373911740375… ┆ {"Content-Type":"application/… │ ├╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┤ │ b8780f2e-eea4-4c8d-9e68-757bd… ┆ http://0731yhwj.com/jjfa/1636… ┆ response ┆ 2025-09-05 12:16:36 UTC ┆ 106231 ┆ text/html ┆ b"HTTP/1.1 200 OK\r\nDate: Fr… ┆ {"Content-Type":"application/… │ ├╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┤ │ 960529dc-840a-42df-b9f3-200c0… ┆ http://0731yhwj.com/jjfa/1636… ┆ metadata ┆ 2025-09-05 12:16:36 UTC ┆ 296 ┆ None ┆ b"fetchTimeMs: 696\r\ncharset… ┆ {"Content-Type":"application/… │ ├╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┤ │ 0657b53a-5691-41ed-bbf6-53d88… ┆ http://1-apple.com.tw/index.c… ┆ request ┆ 2025-09-05 11:49:26 UTC ┆ 401 ┆ None ┆ b"GET /index.cfm?Fuseaction=M… ┆ {"Content-Type":"application/… │ ╰────────────────────────────────┴────────────────────────────────┴───────────┴─────────────────────────────────────────┴────────────────┴──────────────────────────────┴────────────────────────────────┴────────────────────────────────╯ (Showing first 8 rows) >>> daft.datasets.common_crawl("CC-MAIN-2025-33", segment="1754151279521.11").limit(1).show() ╭────────────────────────────────┬─────────────────┬───────────┬─────────────────────────────────────────┬────────────────┬──────────────────────────────┬────────────────────────────────┬────────────────────────────────╮ │ WARC-Record-ID ┆ WARC-Target-URI ┆ WARC-Type ┆ WARC-Date ┆ Content-Length ┆ WARC-Identified-Payload-Type ┆ warc_content ┆ warc_headers │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ Utf8 ┆ Utf8 ┆ Utf8 ┆ Timestamp(Nanoseconds, Some("Etc/UTC")) ┆ Int64 ┆ Utf8 ┆ Binary ┆ Utf8 │ ╞════════════════════════════════╪═════════════════╪═══════════╪═════════════════════════════════════════╪════════════════╪══════════════════════════════╪════════════════════════════════╪════════════════════════════════╡ │ 526c37b2-f535-4015-b8dd-bfa8e… ┆ None ┆ warcinfo ┆ 2025-08-02 22:09:07 UTC ┆ 489 ┆ None ┆ b"isPartOf: CC-MAIN-2025-33\r… ┆ {"Content-Type":"application/… │ ╰────────────────────────────────┴─────────────────┴───────────┴─────────────────────────────────────────┴────────────────┴──────────────────────────────┴────────────────────────────────┴────────────────────────────────╯ (Showing first 1 of 1 rows) >>> daft.datasets.common_crawl("CC-MAIN-2025-33", segment="1754151279521.11", content="metadata").limit(1).show() ╭────────────────────────────────┬─────────────────┬───────────┬─────────────────────────────────────────┬────────────────┬──────────────────────────────┬────────────────────────────────┬────────────────────────────────╮ │ WARC-Record-ID ┆ WARC-Target-URI ┆ WARC-Type ┆ WARC-Date ┆ Content-Length ┆ WARC-Identified-Payload-Type ┆ warc_content ┆ warc_headers │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ Utf8 ┆ Utf8 ┆ Utf8 ┆ Timestamp(Nanoseconds, Some("Etc/UTC")) ┆ Int64 ┆ Utf8 ┆ Binary ┆ Utf8 │ ╞════════════════════════════════╪═════════════════╪═══════════╪═════════════════════════════════════════╪════════════════╪══════════════════════════════╪════════════════════════════════╪════════════════════════════════╡ │ cfae7e3e-02b7-4e13-b94e-8e9dd… ┆ None ┆ warcinfo ┆ 2025-08-16 01:03:20 UTC ┆ 277 ┆ None ┆ b"Software-Info: ia-web-commo… ┆ {"Content-Type":"application/… │ ╰────────────────────────────────┴─────────────────┴───────────┴─────────────────────────────────────────┴────────────────┴──────────────────────────────┴────────────────────────────────┴────────────────────────────────╯ (Showing first 1 of 1 rows) >>> daft.datasets.common_crawl("CC-MAIN-2025-33", segment="1754151279521.11", content="wet").limit(1).show() ╭────────────────────────────────┬─────────────────┬───────────┬─────────────────────────────────────────┬────────────────┬──────────────────────────────┬────────────────────────────────┬────────────────────────────────╮ │ WARC-Record-ID ┆ WARC-Target-URI ┆ WARC-Type ┆ WARC-Date ┆ Content-Length ┆ WARC-Identified-Payload-Type ┆ warc_content ┆ warc_headers │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ Utf8 ┆ Utf8 ┆ Utf8 ┆ Timestamp(Nanoseconds, Some("Etc/UTC")) ┆ Int64 ┆ Utf8 ┆ Binary ┆ Utf8 │ ╞════════════════════════════════╪═════════════════╪═══════════╪═════════════════════════════════════════╪════════════════╪══════════════════════════════╪════════════════════════════════╪════════════════════════════════╡ │ 0cb039e8-d357-485f-95dd-cdfdb… ┆ None ┆ warcinfo ┆ 2025-08-16 01:03:20 UTC ┆ 370 ┆ None ┆ b"Software-Info: ia-web-commo… ┆ {"Content-Type":"application/… │ ╰────────────────────────────────┴─────────────────┴───────────┴─────────────────────────────────────────┴────────────────┴──────────────────────────────┴────────────────────────────────┴────────────────────────────────╯ (Showing first 1 of 1 rows) >>> daft.datasets.common_crawl("CC-MAIN-2025-33", num_files=1).limit(1).show() ╭────────────────────────────────┬─────────────────┬───────────┬─────────────────────────────────────────┬────────────────┬──────────────────────────────┬────────────────────────────────┬────────────────────────────────╮ │ WARC-Record-ID ┆ WARC-Target-URI ┆ WARC-Type ┆ WARC-Date ┆ Content-Length ┆ WARC-Identified-Payload-Type ┆ warc_content ┆ warc_headers │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ Utf8 ┆ Utf8 ┆ Utf8 ┆ Timestamp(Nanoseconds, Some("Etc/UTC")) ┆ Int64 ┆ Utf8 ┆ Binary ┆ Utf8 │ ╞════════════════════════════════╪═════════════════╪═══════════╪═════════════════════════════════════════╪════════════════╪══════════════════════════════╪════════════════════════════════╪════════════════════════════════╡ │ 526c37b2-f535-4015-b8dd-bfa8e… ┆ None ┆ warcinfo ┆ 2025-08-02 22:09:07 UTC ┆ 489 ┆ None ┆ b"isPartOf: CC-MAIN-2025-33\r… ┆ {"Content-Type":"application/… │ ╰────────────────────────────────┴─────────────────┴───────────┴─────────────────────────────────────────┴────────────────┴──────────────────────────────┴────────────────────────────────┴────────────────────────────────╯ (Showing first 1 of 1 rows) See https://github.com/Eventual-Inc/Daft/discussions/5248 for more discussions and followups. | 1 年前 | |
feat: Add WARC reader (#3871) Adds a reader for .warc and .warc.gz files. Currently optimized for reading from S3. Some numbers: - Downloading a single common crawl file from S3, e.g. s3://commoncrawl/crawl-data/CC-MAIN-2018-17/segments/1524125937193.1/warc/CC-MAIN-20180420081400-20180420101400-00000.warc.gz, takes ~16.5s. - Gunzipping this file and processing it with fastwarc takes ~33s. - With swordfish, collecting that same file as a daft dataframe takes ~20s on an m7g.4xlarge instance. - With swordfish, collecting 10 common crawl files take ~3min 12s. - With swordfish, processing (read then sum on content length) 1 common crawl file takes ~20s. - With swordfish, processing 2 common crawl files still takes ~20s. - With swordfish, processing 10 common crawl files takes ~40s. This is because we've set the max number of parallel reads to 8. So 10 scan tasks take 2x20s to read. If we increase the max number of parallel reads to 10, the runtime drops to ~30s. - With ray, collecting 1 file takes ~25s. - **Unfortunately, with ray, collecting 10 files caused the instance to become unresponsive.** Followup work: - Extracting fields from the warc_headers json is not very fast. We can do better here by allowing users to specify the metadata headers that they want to extract. --------- Co-authored-by: Sammy Sidhu <sammy.sidhu@gmail.com> Co-authored-by: Colin Ho <colinho@Colins-MBP.localdomain> | 1 年前 | |
feat: Add WARC reader (#3871) Adds a reader for .warc and .warc.gz files. Currently optimized for reading from S3. Some numbers: - Downloading a single common crawl file from S3, e.g. s3://commoncrawl/crawl-data/CC-MAIN-2018-17/segments/1524125937193.1/warc/CC-MAIN-20180420081400-20180420101400-00000.warc.gz, takes ~16.5s. - Gunzipping this file and processing it with fastwarc takes ~33s. - With swordfish, collecting that same file as a daft dataframe takes ~20s on an m7g.4xlarge instance. - With swordfish, collecting 10 common crawl files take ~3min 12s. - With swordfish, processing (read then sum on content length) 1 common crawl file takes ~20s. - With swordfish, processing 2 common crawl files still takes ~20s. - With swordfish, processing 10 common crawl files takes ~40s. This is because we've set the max number of parallel reads to 8. So 10 scan tasks take 2x20s to read. If we increase the max number of parallel reads to 10, the runtime drops to ~30s. - With ray, collecting 1 file takes ~25s. - **Unfortunately, with ray, collecting 10 files caused the instance to become unresponsive.** Followup work: - Extracting fields from the warc_headers json is not very fast. We can do better here by allowing users to specify the metadata headers that they want to extract. --------- Co-authored-by: Sammy Sidhu <sammy.sidhu@gmail.com> Co-authored-by: Colin Ho <colinho@Colins-MBP.localdomain> | 1 年前 | |
[FEAT]: SQL read_csv (#3255) Add the functionality to read from csv file in a sql query --------- Co-authored-by: Itzhak Stern <itzhaks@ubup039.me-corp.lan> | 1 年前 | |
[BUG] Azure and Iceberg read and write fixes (#2349) In this PR: - pyarrow.dataset.write_dataset does not properly write Parquet metadata in version 12.0.0, set the requirements for it to be >=12.0.1 - Azure fsspec filesystem now initialized IOConfig values - Azure URIs that look like PROTOCOL://account.dfs.core.windows.net/container/path-part/file now properly parsed, URI parsing also cleaned up and unified - fixed small discrepancies for AzureConfig in daft.pyi - Added a public test Iceberg table on Azure, a SQLite catalog that points to the table, and a test for those tables. - More tests should be written - #2348 Should resolve #2005 | 2 年前 | |
feat: audio file subtype (#5602) ## Changes Made Creates new AudioFile class that's mostly a wrapper around soundfile library. provides common methods for interacting with audio files such as - metadata - resample - to_numpy Also supports dataframe expressions - daft.functions.audio_file - daft.functions.resample - daft.functions.audio_metadata Still need to add tests and a few more methods. Will follow up in a separate PR some functionality for audio extraction from video files. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly | 10 个月前 | |
feat: add a new subtype of file for video ops (#5346) ## Changes Made ### Summary - adds a mediatype argument for file datatype. DataType.file(MediaType.unknown() |MediaType.video()) - adds new uv/pip category: daft[video] - adds a new VideoFile subclass that is used in place of daft.File (for udfs) if it's known to be a video file. Examples py df = daft.from_glob_paths("**/*.mp4) df = df.select(daft.functions.video_file(df["path"]).alias("video")) # can also enable runtime validation of the video files df = df.select(daft.functions.video_file(df["path"], verify=True).alias("video")) # get the metadata from the files df = df.select("*", daft.functions.video_metadata(df["video"]).unnest()) # get the keyframes df = df.select("*", daft.functions.video_keyframes(df["video"])) Can also use these in udfs py from daft.file.typing import VideoMetadata df = daft.from_glob_paths("**/*.mp4) df = df.select(daft.functions.video_file(df["path"]).alias("video")) @daft.func def metadata(f: daft.VideoFile)->VideoMetadata: return f.metadata() @daft.func def keyframes(f: daft.VideoFile)->list[PIL.Image.Image] return list(f.keyframes()) df = df.select(metadata(df["video"]), keyframes(df["video"])) can also use them as standalone python objects py file = daft.VideoFile("path/to/video.mp4") keyframes = list(file.keyframes()) Additional notes. I had to get rid of the common-file and just move all of that in to daft-core ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly | 11 个月前 | |
[PERF] Predicate Pushdown for CSV Reader (#1724) * Enables predicate and early termination in the CSV reader * Also refactors our FromArrow to take in a FieldRef | 2 年前 | |
[PERF] Json Predicate Pushdown while reading (#1727) | 2 年前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 1 年前 | ||
| 1 年前 | ||
| 2 年前 | ||
| 2 年前 | ||
| 3 年前 | ||
| 1 年前 | ||
| 1 年前 | ||
| 1 年前 | ||
| 1 年前 | ||
| 2 年前 | ||
| 10 个月前 | ||
| 11 个月前 | ||
| 2 年前 | ||
| 2 年前 |