| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
chore: Remove the old Ray Runner (#5375) ## Changes Made 🎉🎂🥳 Can finally delete it, Flotilla supports all necessary features. Additional Related Features Removed: * DataFrame.num_partitions: This is computed using the old Ray runner's planner. We could use the new Ray runner, but it seems kind of unnecessary * Context Settings: --------- Co-authored-by: Colin Ho <colin.ho99@gmail.com> | 9 个月前 | |
feat(ai): Add HTTP URL direct passthrough for images and videos in prompt function (#6182) ## Changes Made This PR adds support for HTTP/HTTPS URL direct passthrough for images and videos in the prompt function. Previously, all media content had to be downloaded and base64-encoded before being sent to the LLM. Now, publicly accessible URLs are passed directly to the model, significantly reducing I/O overhead and memory consumption. <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 7 个月前 | |
feat: audio file subtype (#5602) ## Changes Made Creates new AudioFile class that's mostly a wrapper around soundfile library. provides common methods for interacting with audio files such as - metadata - resample - to_numpy Also supports dataframe expressions - daft.functions.audio_file - daft.functions.resample - daft.functions.audio_metadata Still need to add tests and a few more methods. Will follow up in a separate PR some functionality for audio extraction from video files. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly | 10 个月前 | |
chore: use dedicated OSS AWS account (#6442) ## Changes Made Update to use new dedicated OSS AWS account: - S3 buckets for GitHub artifacts, and public datasets - CloudFront distribution Requires: - Update ACTIONS_AWS_IAM_ROLE GitHub secret ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 6 个月前 | |
feat!: add iceberg table support with the gravitino catalog (#6509) ## Changes Made This refactors the Gravitino catalog implementation to match our Glue implementation. It looks at table information to determine which concrete Table implementation to use. This also refactors the Catalog creation a little bit to once again match Glue. I'm not proud of the Catalog.from_* thing with the objects, and I like the pyiceberg choice of the top-level load_ methods, hence why later implementations like Glue use that pattern, so I've updated Gravitino to follow this. This PR also properly internalizes/encapsulates the client, I wasn't interested in having our own Gravitino client vs. using the python one, but the customer required this -- so I just hid the client and updated the docs accordingly. Also a tiny fix on an old typo of mine from the video reader: read_glob_paths/from_glob_paths. ## Related Issues - Closes #6318 - Closes #6319 | 5 个月前 | |
fix: close connections in io tests and address other warnings (#6470) ## Changes Made * Iceberg tests weren't closing the sqlite connection which made our CI logs flooded with warnings. * Several other io tests had the same issue. * Also fixes some additional warnings which are making CI logs noisy. ## Related Issues N/A | 6 个月前 | |
chore: Move Flight Server to Rust (#6519) ## Changes Made Provide in Rust so we have direct access to it for optimizations | 5 个月前 | |
feat: Explicit AWS vs. HTTP mode for common crawl dataset (#5379) Adds a new required argument to daft.datasets.common_crawl: in_aws: bool. This **must** be set to True when running in AWS and False when running outside of AWS. This allows Daft to select the most optimal download strategy for CC data. Added a notice about this to the docstring. Refactors the existing mocked unit tests for this by making the tests patch the appropriate _get_{s3,http}_manifest_path using the value of in_aws. Adds in_aws as a pytest parameter and parameterizes each test on True and False. Updates the Common Crawl documentation to mention the new required in_aws parameter. Adds a new section discussing the new HTTP download mode and provides an example. | 11 个月前 | |
fix: short-circuit evaluation for coalesce (#6525) ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> Fixes #4069 by adding proper short-circuit behavior to coalesce. - Added Expr::Coalesce with short-circuit logic - Implemented early-exit in record batch evaluation - Preserved semantics during UDF optimization - Added short-circuit tests ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> Closes #4069 --------- Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> | 5 个月前 | |
feat: adds support for daft native extensions via stable C ABI (#6301) ## About This PR creates a native extension framework for Daft and is modeled after Postgres. Extension authors ship a normal pip-installable python package with a bundled dylib (instructions in tutorial). Users import python functions to get full IDE support and the native implementations are linked by loading the extension into a session. ## Design **Crates** - daft-ext - Single dependency for extension authors (published). - daft-ext-abi - Stable #[repr(C)] types defining the ABI contract. - daft-ext-core - Public Rust SDK types (DaftScalarFunction, traits, error types). - daft-ext-macros - Proc macro (#[daft_extension]). - daft-ext-internal - Host-side adapters (ScalarFunctionHandle, module loader). Not published. **Functions** ScalarUDF is Daft's internal trait for scalar functions which cannot cross a dlopen boundary because it uses Rust trait objects with unstable ABI layouts - so FFI_ScalarFunction in daft-ext-abi is the stable C ABI version of this. ScalarFunctionHandle in daft-ext-internal wraps a FFI_ScalarFunction and implements both ScalarUDF and ScalarFunctionFactory. Data crosses the extension boundary using the Arrow C Data Interface (FFI_ArrowArray + FFI_ArrowSchema) and is zero-copy. The extension FFI uses arrow::ffi directly, not the common-arrow-ffi crate which is coupled to PyO3. **Installation** We load shared libs once into the process when someone calls .load_extension on the session. The top-level daft.load_extension will load the given extension module into the active session (from context). All defined functions are scoped to the session in which the extension is loaded. ## Changes - Creates the daft-ext-* crates defined above. - Defines the stable C ABI which minimally wraps arrow ffi. - Creates the daft-ext-internal to bridge host->abi. - Creates the daft-ext-core to bridge extension->abi. - Implements internal Daft traits using the daft-ext-internal types. - Adds a dvector example extension which is a pgvector clone. - Adds a hello example extension with a tutorial document. - Extends the session to support native function registration. - Adds the get_function method for session-backed function resolution mirroring the SQL implementation. Can add more clarity upon reviews. ## Examples - [Hello, World](https://github.com/Eventual-Inc/Daft/tree/564548a93fa99180e4e17f51049e5db60a5949cf/examples/hello) - [dvector (pgvector)](https://github.com/Eventual-Inc/Daft/tree/564548a93fa99180e4e17f51049e5db60a5949cf/examples/dvector) python import daft # Step 1. Import your extension module import hello # Step 2. Load the extension into the current daft session daft.load_extension(hello) # Step 3. Use in your dataframe! df = daft.from_pydict({"name": ["John", "Paul"]}) df = df.select(hello.greet(df["name"])) See the actual [greet rust implementation](https://github.com/Eventual-Inc/Daft/blob/564548a93fa99180e4e17f51049e5db60a5949cf/examples/hello/src/lib.rs). ## Guide See [Daft Extension Guide](https://github.com/Eventual-Inc/Daft/blob/564548a93fa99180e4e17f51049e5db60a5949cf/docs/extensions/index.md). | 6 个月前 | |
fix(video): add missing pillow dependency for video keyframes (#6355) Based on @everettVT's branch. video_keyframes() fails with a confusing DaftError::ValueError Need at least 1 series to perform concat when Pillow is not installed. This primarily affects Python 3.12+ users because the daft[video] extra didn't include pillow as a dependency, and Pillow isn't transitively installed on those versions. The root cause is that VideoFile.keyframes() calls frame.to_image() (which requires PIL) without checking for its availability first. When every row fails, the errors get swallowed and ListArray::try_from hits an empty concat instead of surfacing the actual ModuleNotFoundError. To fix this, we add pillow==12.1.1 to the video extra in pyproject.toml (matching the pin used by the google, openai, and transformers extras) and add an early import check in VideoFile.keyframes() that raises a clear ImportError directing users to pip install daft[video]. Fixes #6064 Co-authored-by: Everett Kleven <ekleven@vangelis.tech> | 6 个月前 | |
feat: Add hex and unhex (#6373) ## Changes Made - Integrated hex encoding/decoding into existing encode/decode functions as a "hex" charset option <!-- Describe what changes were made and why. Include implementation details if necessary. --> ## Related Issues #3793 <!-- Link to related GitHub issues, e.g., "Closes #123" --> --------- Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> Co-authored-by: R. C. Howell <5731503+rchowell@users.noreply.github.com> | 6 个月前 | |
feat!: add iceberg table support with the gravitino catalog (#6509) ## Changes Made This refactors the Gravitino catalog implementation to match our Glue implementation. It looks at table information to determine which concrete Table implementation to use. This also refactors the Catalog creation a little bit to once again match Glue. I'm not proud of the Catalog.from_* thing with the objects, and I like the pyiceberg choice of the top-level load_ methods, hence why later implementations like Glue use that pattern, so I've updated Gravitino to follow this. This PR also properly internalizes/encapsulates the client, I wasn't interested in having our own Gravitino client vs. using the python one, but the customer required this -- so I just hid the client and updated the docs accordingly. Also a tiny fix on an old typo of mine from the video reader: read_glob_paths/from_glob_paths. ## Related Issues - Closes #6318 - Closes #6319 | 5 个月前 | |
chore(io): replace python source shim on the rust side (#6556) ## Motivation This removes the python ScanOperator shim in favor of a ScanOperator implementation in Rust which is backed by either a Python *or* Rust data source. This allows us to move the shim one step lower, and now we actually have Rust traits + python bridges for the DataSource and DataSourceTask interfaces. This is an incremental step towards full streaming sources and consolidating the various scanning types and implementations to this interface. ## Changes Made - Removes python _DataSourceShim - Cleanup visibility modifiers and conversion methods for daft-catalog Catalog, Table, Provider - Adds native factory methods for native data source tasks, which are a work-in-progress and just today's scan tasks. - We have a Rust PyDataSourceWrapper trait which is able to implement our existing ScanOperator backed by a Python DataSource | 5 个月前 | |
chore: use dedicated OSS AWS account (#6442) ## Changes Made Update to use new dedicated OSS AWS account: - S3 buckets for GitHub artifacts, and public datasets - CloudFront distribution Requires: - Update ACTIONS_AWS_IAM_ROLE GitHub secret ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 6 个月前 | |
fix(observability): Truncate progress bar names by characters, not bytes (#6180) ## Changes Made Progress bar pipeline name truncation used byte-level slicing which panics on mulit-byte UTF-8 characters. Fix this by truncating by character length instead. ## Related Issues See: https://github.com/Eventual-Inc/Daft/actions/runs/21921434809 | 7 个月前 | |
test: Filter null bytes from generated column names in property-based tests (#6213) ## Summary - Filters embedded null bytes ( \0) from Hypothesis-generated column names in property-based tests - Fixes a panic in arrow-rs FFI layer (FFI_ArrowSchema::with_name calls CString::new(name).unwrap()) triggered when Ray serializes columns with null bytes in their names - The failing CI run: https://github.com/Eventual-Inc/Daft/actions/runs/22042235355 ## Root Cause The Hypothesis text() strategy can generate strings containing null bytes. When such a string is used as a column name and the test runs on the Ray runner, the serialization path (Series.__reduce__ → to_arrow() → Arrow C Data Interface) hits a panic in arrow-rs v57.2.0 at arrow-schema/src/ffi.rs:165: rust // FFI_ArrowSchema::with_name() - panics instead of returning error self.name = CString::new(name).unwrap().into_raw(); Null bytes in column names are inherently invalid for the Arrow C Data Interface (which uses null-terminated C strings), so filtering them from generated test data is the correct fix. The upstream arrow-rs bug (.unwrap() instead of error propagation) still exists in v57.3.0. Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> | 7 个月前 | |
feat: support to_ray_dataset() from the native runner (#6486) DataFrame.to_ray_dataset() previously required the Ray runner, raising a ValueError when called from the native runner. This forced users who run Daft compute on the native runner (for performance) but need Ray Datasets for downstream ML pipelines to do manual Arrow/pylist round-trips. The restriction was artificial. When using the native runner, partitions can be converted to Arrow tables and passed directly to Ray viaray.data.from_arrow(). --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> | 6 个月前 | |
chore: Collection of random Swordfish cleanups (#6454) ## Changes Made A bunch of random cleanup / chores that I did in Swordfish as I was cleaning up the PR to add real progress bars. Rather than clutter that PR, I split it out into this one. Some things in particular: * Made RuntimeStats an associated type in local PipelineNodes * Changed the RuntimeStats trait a bit in response * Made operators return MicroPartitions directly instead of Arc<MicroPartition> & making MicroPartitions cloneable * I think they weren't originally because of the Unloaded option, but I don't think that's necessary anymore. Feel free to lmk if you disagree * Removed TimedFutures and instead recycled the timers we use for dynamic batching. This actually ended up fixing a bug we saw with the timers. | 6 个月前 | |
feat: support various image hash functions (aHash, dHash, pHash, wHash, crop-resistant hash) for deduplication (#6338) Implements 5 image hashing algorithms for image deduplication: aHash (Average Hash): Compares each pixel to the mean intensity dHash (Difference Hash): Compares adjacent pixel differences pHash (Perceptual Hash): DCT-based frequency domain hashing wHash (Wavelet Hash): Haar wavelet transform-based hashing Crop-resistant Hash: Robust against cropping transformations ## Changes Made Added hash methods to CowImage with helper functions (dct2d_32x32, haar_2d) Added ImageHashOps trait and implementations for ImageArray / FixedShapeImageArray Added series-level functions and ScalarUDF registrations Exposed Python APIs: image_ahash, image_dhash, image_phash, image_whash, image_crop_resistant_hash Added corresponding methods to SeriesImageNamespace Added tests covering all 5 algorithms Fixed-size hashes return FixedSizeBinary(8), crop-resistant hash returns variable-length Binary. ## Related Issues Closes #4889 | 5 个月前 | |
fix: short-circuit evaluation for coalesce (#6525) ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> Fixes #4069 by adding proper short-circuit behavior to coalesce. - Added Expr::Coalesce with short-circuit logic - Implemented early-exit in record batch evaluation - Preserved semantics during UDF optimization - Added short-circuit tests ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> Closes #4069 --------- Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> | 5 个月前 | |
feat(swordfish): plan caching (#6278) ## Summary Introduces fingerprint-based plan caching so multiple executions of the same logical plan share a single pipeline, rather than building a new one each time. ### Key changes - ** PipelineMessage enum** — Replaces raw MicroPartition in all pipeline channels with Morsel { input_id, partition } and Flush(input_id), allowing a single pipeline to multiplex data from multiple logical inputs - **NativeExecutor becomes stateful** — Maintains a HashMap<u64, PlanState> keyed by plan fingerprint; repeated calls to run() with the same fingerprint reuse the existing pipeline - **MessageRouter** — Routes pipeline output to per-input_id unbounded channels so each caller gets only its own results - **try_finish() lifecycle API** — Callers signal input completion and collect per-input stats; pipeline is torn down when the last input finishes - **Per-input runtime stats** — RuntimeStatsManager tracks (NodeID, InputId) pairs with merge support for aggregated views - **next_event helper** — Shared tokio::select! loop abstraction used by blocking sinks, streaming sinks, intermediate ops, and joins - **Simplified BlockingSink trait** — Removed BlockingSinkFinalizeOutput::HasMoreOutput (was dead code); finalize now returns Vec<MicroPartition> directly; make_state takes InputId - **Simplified StreamingSink trait** — Removed StreamingSinkOutput::HasMoreOutput variant - **All node types updated** — Sources, intermediate ops, blocking sinks, streaming sinks, joins, and concat all handle Morsel/Flush per-input routing - **Python integration** — native_executor.py and flotilla.py call try_finish() for lifecycle management ### Files changed ~49 files across daft-local-execution (pipeline, sources, sinks, joins, stats), daft-distributed, common-metrics, and Python runners. ## Test plan - [ ] Full test suite with DAFT_RUNNER=native - [ ] Full test suite with DAFT_RUNNER=ray - [ ] Concurrent query execution sharing the same plan fingerprint - [ ] Verify per-input stats are correctly reported and merged - [ ] TPC-H benchmarks to check for performance regressions 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: R. C. Howell <5731503+rchowell@users.noreply.github.com> | 5 个月前 | |
chore: Remove expression namespaces (#5619) ## Changes Made Bye bye ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly | 9 个月前 | |
feat(swordfish): plan caching (#6278) ## Summary Introduces fingerprint-based plan caching so multiple executions of the same logical plan share a single pipeline, rather than building a new one each time. ### Key changes - ** PipelineMessage enum** — Replaces raw MicroPartition in all pipeline channels with Morsel { input_id, partition } and Flush(input_id), allowing a single pipeline to multiplex data from multiple logical inputs - **NativeExecutor becomes stateful** — Maintains a HashMap<u64, PlanState> keyed by plan fingerprint; repeated calls to run() with the same fingerprint reuse the existing pipeline - **MessageRouter** — Routes pipeline output to per-input_id unbounded channels so each caller gets only its own results - **try_finish() lifecycle API** — Callers signal input completion and collect per-input stats; pipeline is torn down when the last input finishes - **Per-input runtime stats** — RuntimeStatsManager tracks (NodeID, InputId) pairs with merge support for aggregated views - **next_event helper** — Shared tokio::select! loop abstraction used by blocking sinks, streaming sinks, intermediate ops, and joins - **Simplified BlockingSink trait** — Removed BlockingSinkFinalizeOutput::HasMoreOutput (was dead code); finalize now returns Vec<MicroPartition> directly; make_state takes InputId - **Simplified StreamingSink trait** — Removed StreamingSinkOutput::HasMoreOutput variant - **All node types updated** — Sources, intermediate ops, blocking sinks, streaming sinks, joins, and concat all handle Morsel/Flush per-input routing - **Python integration** — native_executor.py and flotilla.py call try_finish() for lifecycle management ### Files changed ~49 files across daft-local-execution (pipeline, sources, sinks, joins, stats), daft-distributed, common-metrics, and Python runners. ## Test plan - [ ] Full test suite with DAFT_RUNNER=native - [ ] Full test suite with DAFT_RUNNER=ray - [ ] Concurrent query execution sharing the same plan fingerprint - [ ] Verify per-input stats are correctly reported and merged - [ ] TPC-H benchmarks to check for performance regressions 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: R. C. Howell <5731503+rchowell@users.noreply.github.com> | 5 个月前 | |
add tests | 4 年前 | |
chore: optimizes slow tests in CI/CD (#6029) ## Results The PR reduced unit-test time (3.13, native, 22.0.0, ubuntu-latest, false) from [33m 18s](https://github.com/Eventual-Inc/Daft/actions/runs/20981828603/job/60308365885?pr=6023) to 13m 44s — saving about 20 mins which is nice. The test_context is still slow, but it has very few tests, a big win is parallelization and taking >5 mins off just by skipping the setup for skipped tests. This also reduced the ray tests from 45 mins to 27 mins, and there should be plenty more savings here. Then disabling windows eliminated a 50+ min action. text ================================================================================ OVERALL SUMMARY ================================================================================ Tests Analyzed: 9620 Total Time: 941.32s Average Time: 0.05s Max Time: 14.70s Min Time: 0.01s It took 188s to run make test on my mac which is nice; it also crushed my machine while it was running 🥲 ## Changes Made * Adds -n auto to pytest for parallelize testing on multiple cores * Skips generating fixtures (expensive?) by not actually creating the tests - we were creating test data _then_ skipping for several thousand join tests taking ~5 mins of setup work just to skip. * Memoizes data generation for the window tests (often the slowest for me) * Reduces the data generation size and faster comparison algo (before was 10,000 data points with a O(n^2) algo, now it's 1,000 data points with an O(nlogn) comparison algo) * Disables the RAY_DASHBOARD for the context setup commands, and other minor things to speedup ray init * Skips windows on PRs (>50 minutes) only enabled on main, worthwhile trade-off because we will still catch windows-specific failures when merging to main which is quite rare to be fair. ## Related Issues N/A | 8 个月前 | |
chore(tests): migrate internal usages of daft.udf to cls/func (#6348) update tests to remove a bunch of deprecation warnings in CI/tests ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 6 个月前 | |
chore: remove _FixEmptyStructArrays workaround (#6478) ## Changes Made Removes the _FixEmptyStructArrays workaround from the codebase. This was legacy code that converted empty struct arrays (pa.struct([])) into single-field placeholder structs (pa.struct({"": pa.null()})) on ingestion, then reversed the transformation on output. Arrow now properly supports empty structs, and the root cause in Daft — StructArray deriving its length from the first child (yielding 0 for zero-field structs) — is fixed directly. **Core fix:** - StructArray::new now derives length from the null buffer when children is empty - Added StructArray::new_empty constructor for zero-field structs with explicit length - StructArray::to_arrow uses arrow::array::StructArray::new_empty_fields for empty structs **Removed placeholder machinery:** - Deleted daft/arrow_utils.py (_FixEmptyStructArrays, remove_empty_struct_placeholders, ensure_* helpers) - Removed all callers in series.py, recordbatch.py, ray_runner.py - Removed remove_empty_struct_placeholders call from Rust FFI path (series.rs) - Removed {String::new() => Literal::Null} placeholder pattern in pydict_to_struct_lit / pytuple_to_struct_lit - Removed placeholder field filters in literal-to-Python conversion, repr, JSON inference, and get_lit ## Related Issues Follow-up to #6378 which kept _FixEmptyStructArrays due to internal limitations that are now resolved. | 6 个月前 | |
chore: Drop Python 3.9 (#5479) ## Changes Made Python 3.9 is EOL starting from Nov 1, so lets drop it and make 3.10 the minimum. Also, start testing ranges 3.10 to 3.13 --------- Co-authored-by: Colin Ho <colin.ho99@gmail.com> | 10 个月前 | |
feat(subscriber): add Event enum and on_event dispatch for subscriber (#6508) ## Changes Made Introduces Event transport enum with OperatorStart/OperatorEnd/Stats variants and a single on_event async method on the Subscriber trait. Ports RuntimeStatsManager to emit all events through on_event, replacing direct per-method calls. Updates dashboard, debug, and python subscribers to implement on_event. | 5 个月前 | |
feat: Add pow expression (#5237) ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> ## Related Issues #4704 <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly (tag @/ccmao1130 for docs review) --------- Co-authored-by: Colin Ho <colin.ho99@gmail.com> | 10 个月前 | |
feat: Better errors when lazy imports fail (#5753) Accessing an attribute on a lazy-imported module will now raise an ImportError with an informative error message, rather than a fairly cryptic AttributeError like the following. | 9 个月前 | |
fix: make it easier to enable different logging levels (#5661) ## Changes Made make it easier to enable different logging levels. Here's a summary of the PR. - Added convenient logger setup functions: setup_debug_logger, setup_info_logger, setup_warn_logger, and setup_error_logger for quickly configuring different log levels. - Centralized log level configuration logic in setup_logger_level, which supports setting the log level, filtering by prefix, and restricting logs to Daft module only. - When setting the log level, all handler filters are cleared before new filters are added, allowing flexible log filtering. - Added validation for log level input to ensure only valid levels are accepted. - After each logger setup, refresh_logger() is called to synchronize the configuration with the Rust backend. ## Related Issues Closes #5651 | 9 个月前 | |
chore(tests): migrate internal usages of daft.udf to cls/func (#6348) update tests to remove a bunch of deprecation warnings in CI/tests ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 6 个月前 | |
feat: add dual-send telemetry to osstelemetry.io (#6540) ## Summary Send telemetry data to osstelemetry.io alongside Scarf for side-by-side comparison during evaluation. ## Changes - ** daft/scarf_telemetry.py**: After the existing Scarf GET request, send the same data to https://osstelemetry.io/data with activity_type (mapped from endpoint name) and package=daft as additional params. The osstelemetry.io call is independent - failures are silently caught and do not affect existing Scarf telemetry. - **docs/telemetry.md**: Updated to mention osstelemetry.io alongside Scarf. - **tests/test_scarf_telemetry.py**: Updated test_scarf_telemetry_basic to verify both the Scarf URL (first call) and the osstelemetry.io URL (second call) using call_args_list instead of call_args. ## Request format The existing Scarf request: GET https://daft.gateway.scarf.sh/daft-import?version=0.3.0&platform=Darwin&python=3.11&arch=arm64 The new osstelemetry.io request: GET https://osstelemetry.io/data?activity_type=daft-import&package=daft&version=0.3.0&platform=Darwin&python=3.11&arch=arm64 For runner events, both include the runner param (e.g. runner=native or runner=ray): GET https://daft.gateway.scarf.sh/daft-runner?version=0.3.0&platform=Darwin&python=3.11&arch=arm64&runner=native GET https://osstelemetry.io/data?activity_type=daft-runner&package=daft&version=0.3.0&platform=Darwin&python=3.11&arch=arm64&runner=native ## Details - Opt-out flags (SCARF_NO_ANALYTICS, DO_NOT_TRACK, DAFT_ANALYTICS_ENABLED) apply to both endpoints - Both calls run in the same daemon thread, so no additional threads are created - Dev builds still skip all telemetry ## Testing Tested locally by: 1. Creating a test branch with the dev-build check commented out and both URLs pointed to localhost servers (port 8001 for Scarf, port 8002 for osstelemetry) 2. Running a simple HTTP server that logs all incoming request paths and query params 3. Running import daft to trigger the import telemetry event Both servers received the expected requests: [SCARF] GET /daft-import version: 0.3.0-dev0 platform: Darwin python: 3.11 arch: arm64 [OSSTELEMETRY] GET /data activity_type: daft-import package: daft version: 0.3.0-dev0 platform: Darwin python: 3.11 arch: arm64 | 5 个月前 | |
refactor: Cleanup Dtype Names (#5400) | 11 个月前 | |
feat: Support pattern filtering for SHOW TABLES (#5423) ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> ### Extended SQL LIKE Pattern Matching in MemoryCatalog - Upgrade sqlparser to extract namespace from query - Translate SQL LIKE patterns to regex (src/daft-catalog/src/pattern.rs) - Supports standard SQL LIKE wildcards: % (zero or more), _ (exactly one), \ (escape) - Updated documentation on SHOW TABLES syntax and pattern behavior ### Testing - Added pattern matching tests to test_sql_show_tables.py and in src/daft-catalog/pattern.rs - cargo test -p daft-catalog pattern:: ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> Closes #4461 Closes #4007 ## Checklist - [x] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly | 9 个月前 | |
chore: use dedicated OSS AWS account (#6442) ## Changes Made Update to use new dedicated OSS AWS account: - S3 buckets for GitHub artifacts, and public datasets - CloudFront distribution Requires: - Update ACTIONS_AWS_IAM_ROLE GitHub secret ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 6 个月前 | |
feat(subscriber): add Event enum and on_event dispatch for subscriber (#6508) ## Changes Made Introduces Event transport enum with OperatorStart/OperatorEnd/Stats variants and a single on_event async method on the Subscriber trait. Ports RuntimeStatsManager to emit all events through on_event, replacing direct per-method calls. Updates dashboard, debug, and python subscribers to implement on_event. | 5 个月前 | |
fix: Optimize the display information of Join nodes in query plan (#5617) ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> Currently, when displaying the query plan containing a "join" node using explain, we can't intuitively distinguish between the left and right tables, especially in the case of a Broadcast Join, which requires the is_swapped field to assist in judgment. This PR mainly unifies the information format and elements of the Join nodes displayed in explain, and explicitly identifies the Broadcaster and Receiver for Broadcast Join. For example: text == Physical Plan == * BroadcastJoin | Type: Inner | Left: Join key = col(1: l_name), Role = Receiver | Right: Join key = col(1: s_name), Role = Broadcaster | Null equals nulls: [false] |\ | * ScanTaskSource: | | Num Scan Tasks = 100 | | Estimated Scan Bytes = 728327 | | Pushdowns: {filter: not(is_null(col(l_name)))} | | Schema: {id#Int64, l_name#Utf8, l_email#Utf8} | | Scan Tasks: [ ... ] | * Project: col(0: id) as right.id, col(1: s_name), col(2: s_email) | Resource request = None | * ScanTaskSource: | Num Scan Tasks = 10 | Estimated Scan Bytes = 67702 | Pushdowns: {filter: not(is_null(col(s_name)))} | Schema: {id#Int64, s_name#Utf8, s_email#Utf8} | Scan Tasks: [ ... ] ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly Signed-off-by: plotor <zhenchao.wang@hotmail.com> | 8 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 9 个月前 | ||
| 7 个月前 | ||
| 10 个月前 | ||
| 6 个月前 | ||
| 5 个月前 | ||
| 6 个月前 | ||
| 5 个月前 | ||
| 11 个月前 | ||
| 5 个月前 | ||
| 6 个月前 | ||
| 6 个月前 | ||
| 6 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 6 个月前 | ||
| 7 个月前 | ||
| 7 个月前 | ||
| 6 个月前 | ||
| 6 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 9 个月前 | ||
| 5 个月前 | ||
| 4 年前 | ||
| 8 个月前 | ||
| 6 个月前 | ||
| 6 个月前 | ||
| 10 个月前 | ||
| 5 个月前 | ||
| 10 个月前 | ||
| 9 个月前 | ||
| 9 个月前 | ||
| 6 个月前 | ||
| 5 个月前 | ||
| 11 个月前 | ||
| 9 个月前 | ||
| 6 个月前 | ||
| 5 个月前 | ||
| 8 个月前 |