| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
chore: use dedicated OSS AWS account (#6442) ## Changes Made Update to use new dedicated OSS AWS account: - S3 buckets for GitHub artifacts, and public datasets - CloudFront distribution Requires: - Update ACTIONS_AWS_IAM_ROLE GitHub secret ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 6 个月前 | |
chore: Improve Common Crawl Benchmark (#6307) ## Changes Made I was using the common crawl benchmark in our repo to test some things, and I noticed a whole bunch of problems. This PR is a small amalgamation of features to fix those problems. * Add back Scan Operator names that was removed during the streaming sources PR (@colin-ho LMK if you're ok with the solution) * Fix WARC byte counting * Fix bug with glob paths in WARC reads | 7 个月前 | |
test(parquet): add benchmarks for nested types, codecs, and filter pushdown (#6285) Adds benchmark infrastructure for measuring parquet decode performance, establishing baselines ahead of the parquet reader migration from arrow2/parquet2 to arrow-rs. ## Changes Made **Rust microbenchmarks** ( src/daft-parquet/benches/parquet_read.rs) using tango-bench: - Metadata parsing throughput - Single-column int64 decode across codecs (snappy, zstd, uncompressed) - Wide table (20 columns, mixed types) - Dictionary-encoded strings - Nested types (list of int64, struct) - Boolean column **Python benchmarks** (benchmarking/parquet/) using pytest-benchmark: - test_filter_pushdown.py: filter predicate selectivity at 1%/10%/50%/90% across int, float, string, and compound predicates. Daft vs PyArrow side-by-side. - test_types_and_codecs.py: nested/complex types (list, struct, map) × compression codecs (snappy, zstd, gzip, uncompressed). Daft vs PyArrow side-by-side. ## Running the benchmarks bash # Rust microbenchmarks (raw arrow-rs parquet crate decode throughput, 100K rows) cargo bench -p daft-parquet -- solo # Python filter pushdown (build in release mode first: maturin develop --release --uv) DAFT_RUNNER=native python -m pytest benchmarking/parquet/test_filter_pushdown.py \ -o "addopts=" --benchmark-enable --benchmark-columns=mean,stddev,rounds -v # Python types and codecs DAFT_RUNNER=native python -m pytest benchmarking/parquet/test_types_and_codecs.py \ -o "addopts=" --benchmark-enable --benchmark-columns=mean,stddev,rounds -v ## Baseline numbers (release mode, M4 Max) ### Rust microbenchmarks (raw arrow-rs parquet crate, 100K rows) These measure the decode floor — raw arrow-rs parquet reader throughput without any Daft framework overhead. | Benchmark | Median | Range | |-----------|--------|-------| | metadata_parse | 1.0 us | 976 ns – 1.2 us | | single_col_int64_snappy | 442 us | 360 – 985 us | | single_col_int64_zstd | 707 us | 676 – 941 us | | single_col_int64_uncompressed | 380 us | 337 – 456 us | | wide_table_20cols_mixed | 26.3 ms | 26.2 – 26.4 ms | | string_dict_encoded | 840 us | 675 us – 1.3 ms | | nested_list_of_int | 5.0 ms | 4.9 – 5.6 ms | | nested_struct | 5.3 ms | 4.9 – 5.6 ms | | boolean_column | 594 us | 551 – 679 us | This is not useful at the moment, but as we migrate to arrow-rs this will show us whether the bottleneck comes from arrow-rs itself or our implementation. ### Filter pushdown (2M rows, 8 row groups) | Group | Daft (ms) | PyArrow (ms) | Ratio | |-------|-----------|-------------|-------| | filter_int sel_1pct | 25.8 | 15.9 | 1.63x | | filter_int sel_10pct | 27.7 | 14.0 | 1.98x | | filter_int sel_50pct | 32.6 | 17.6 | 1.86x | | filter_int sel_90pct | 36.0 | 16.3 | 2.21x | | filter_float sel_1pct | 25.4 | 13.8 | 1.84x | | filter_float sel_10pct | 27.5 | 14.0 | 1.97x | | filter_float sel_50pct | 34.1 | 17.4 | 1.96x | | filter_float sel_90pct | 35.4 | 15.9 | 2.23x | | filter_str sel_1pct | 26.6 | 14.3 | 1.86x | | filter_str sel_10pct | 28.0 | 14.9 | 1.88x | | filter_str sel_50pct | 33.6 | 18.2 | 1.84x | | filter_str sel_90pct | 37.4 | 16.9 | 2.22x | | filter_compound sel_1pct | 24.7 | 13.0 | 1.90x | | filter_compound sel_10pct | 25.9 | 14.7 | 1.75x | | filter_compound sel_50pct | 28.9 | 17.6 | 1.64x | | filter_compound sel_90pct | 37.6 | 19.2 | 1.95x | | no_filter baseline | 24.0 | 12.5 | 1.91x | Daft is ~1.6–2.2x slower than raw PyArrow on local reads. The fixed overhead (~12ms at sel_1pct vs ~24ms no-filter baseline) is framework cost (query planning, task dispatch) that gets amortized on remote/larger reads. Note that PyArrow uses arrow-cpp (C++ implementation), not arrow-rs, so this comparison reflects different decoder implementations. ### Types and codecs (500K rows, 4 row groups) | Group | Daft (ms) | PyArrow (ms) | Ratio | |-------|-----------|-------------|-------| | list_int64 none | 20.9 | 6.2 | 3.39x | | list_int64 snappy | 25.7 | 9.3 | 2.75x | | list_int64 zstd | 27.9 | 12.4 | 2.25x | | list_int64 gzip | 29.4 | 17.4 | 1.69x | | struct none | 11.9 | 2.6 | 4.53x | | struct snappy | 12.1 | 3.0 | 4.04x | | struct zstd | 12.8 | 3.9 | 3.31x | | struct gzip | 14.9 | 6.2 | 2.41x | | map none | 37.0 | 8.7 | 4.27x | | map snappy | 38.9 | 9.6 | 4.04x | | map zstd | 41.5 | 12.0 | 3.45x | | map gzip | 44.7 | 16.9 | 2.65x | Nested types show larger gaps (3–4.5x for struct/map with fast codecs). The ratio narrows with heavier codecs (gzip) as decompression dominates. See #5741 | 7 个月前 | |
revert!: "revert: Temporarily revert "Remove deprecated APIs for 0.6" (#5084) ## Changes Made Unrevert https://github.com/Eventual-Inc/Daft/pull/5050 now that we've cut a new v0.5 | 1 年前 | |
feat: add Flight shuffle to Flotilla (#6123) Add support for peer-to-peer network transfers for shuffles using the Arrow Flight protocol in the distributed runner. Follow ups: * Add testing infrastructure for distributed TPCH with and without flight shuffle * Support pre-shuffle merge with flight * Refactor Ray-based shuffles to be structured more like Flight * Make shuffle interface more generic | 7 个月前 | |
feat: experimental vllm provider (#5443) ## Changes Made Adds the experimental VLLMPrefixCachedProvider for daft.functions.ai.prompt. Does async batching and prefix routing. When using the VLLMPrefixCachedProvider, prompt will create a VLLMExpr instead of a UDF, which Daft will turn into a custom VLLM operator. This operator is implemented as a streaming sink, and I had to make some minor changes to our streaming sink APIs to make the async batching mechanism work. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly | 11 个月前 | |
Splits TPCH benchmarking and unit tests (#202) * Refactors common data generation utilities into a separate module in benchmarking/tpch/data_generation.py * Refactors TPC-H Daft answers into a separate module in benchmarking/tpch/answers.py * Creates an entrypoint for benchmarking in benchmarking/tpch/__main__.py * Creates an entrypoint for data generation in benchmarking/tpch/data_generation.py * Shares code between benchmarking entrypoint and TPC-H unit tests | 3 年前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 6 个月前 | ||
| 7 个月前 | ||
| 7 个月前 | ||
| 1 年前 | ||
| 7 个月前 | ||
| 11 个月前 | ||
| 3 年前 |