| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 7 个月前 | ||
| 6 个月前 | ||
feat!: add iceberg table support with the gravitino catalog (#6509) ## Changes Made This refactors the Gravitino catalog implementation to match our Glue implementation. It looks at table information to determine which concrete Table implementation to use. This also refactors the Catalog creation a little bit to once again match Glue. I'm not proud of the Catalog.from_* thing with the objects, and I like the pyiceberg choice of the top-level load_ methods, hence why later implementations like Glue use that pattern, so I've updated Gravitino to follow this. This PR also properly internalizes/encapsulates the client, I wasn't interested in having our own Gravitino client vs. using the python one, but the customer required this -- so I just hid the client and updated the docs accordingly. Also a tiny fix on an old typo of mine from the video reader: read_glob_paths/from_glob_paths. ## Related Issues - Closes #6318 - Closes #6319 | 6 个月前 | |
fix: close connections in io tests and address other warnings (#6470) ## Changes Made * Iceberg tests weren't closing the sqlite connection which made our CI logs flooded with warnings. * Several other io tests had the same issue. * Also fixes some additional warnings which are making CI logs noisy. ## Related Issues N/A | 6 个月前 | |
chore: use dedicated OSS AWS account (#6442) ## Changes Made Update to use new dedicated OSS AWS account: - S3 buckets for GitHub artifacts, and public datasets - CloudFront distribution Requires: - Update ACTIONS_AWS_IAM_ROLE GitHub secret ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 6 个月前 | |
chore(tests): migrate internal usages of daft.udf to cls/func (#6348) update tests to remove a bunch of deprecation warnings in CI/tests ## Changes Made <!-- Describe what changes were made and why. Include implementation details if necessary. --> ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 6 个月前 | |
feat(io): implement write_sql with SQLDataSink and explicit dtype support (#5979) ## Summary This PR introduces DataFrame.write_sql() , enabling users to write Daft DataFrames to SQL databases (e.g., PostgreSQL, SQLite) via SQLAlchemy. It implements a robust, distributed SQLDataSink that: 1. Handles Distributed Writes : Uses the DataSink pattern to manage driver-side table initialization and worker-side parallel writes. 2. Supports Explicit Types : Exposes a dtype parameter to allow users to override default type inference with specific SQLAlchemy types (addressing type verification concerns). 3. Ensures Connection Safety : Manages connection lifecycles properly across distributed workers to avoid socket serialization issues. ## Key Changes ### 1. New Public API: DataFrame.write_sql - Location : daft/dataframe/dataframe.py - Signature : def write_sql( self, table_name: str, conn: str | Callable[[], "Connection"], write_mode: Literal["append", "overwrite", "fail"] = "append", chunk_size: int | None = None, dtype: dict[str, Any] | None = None # <--- NEW: Explicit type control ) -> DataFrame: ... - Behavior : Delegates to SQLDataSink and returns a DataFrame with write metrics ( total_written_rows , total_written_bytes ). ### 2. Internal Implementation: SQLDataSink - Location : daft/io/_sql.py - Architecture : - start() (Driver) : Handles write_mode logic. - fail : Checks existence, raises error if table exists, creates table schema if not. - overwrite : Replaces table with new schema. - append : Creates table if not exists, ensuring schema readiness for workers. - write() (Workers) : - Creates independent SQLAlchemy engines/connections per task (no pickling of connections). - Converts MicroPartitions to Pandas. - Uses pd.to_sql with the user-provided dtype to write data efficiently. - Ensures proper resource cleanup ( engine.dispose() ). ### 3. Tests - Location : tests/integration/sql/test_write_sql.py - Coverage : - Sources : Verified with PyDict, CSV, and JSON sources. - Modes : Comprehensive tests for append , overwrite , and fail modes. - Type Verification : Added specific tests ( test_write_sql_dtype_basic_types ) that use sqlalchemy.inspect to verify that columns are created with the correct SQL types when dtype is provided. - Connection Factory : Verified support for passing a connection factory function (crucial for pickling compatibility). ## Addressing Previous Concerns (Type Verification) This implementation addresses concerns about type safety (raised in previous discussions) by: 1. Leveraging Pandas' mature type inference for standard types. 2. Providing the dtype "escape hatch" for complex scenarios, giving users full control over the target schema definition. 3. Including integration tests that explicitly verify schema creation correctness. ## Checklist - I have added comprehensive unit/integration tests. - I have updated the documentation (docstrings included). - I have verified that connections are properly closed and disposed of. Thank you very much for the idea provided by https://github.com/Eventual-Inc/Daft/pull/5471 ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 7 个月前 | |
| 6 个月前 | ||
Initializes working rust-main branch * Old tests are migrated to tests-legacy/ - only new, working tests are left in tests/ * Tests that are failing in tests/ are skipped and tagged with [RUST-INT] * Small fixes made for the code to type-check - removal of udf.py and stubbing out of DataFrame.explode() code | 3 年前 | |
[CHORE] Add TPC-H questions 11-22 to benchmarks (currently skipped) (#2299) | 2 年前 | |
chore(observability): Split dashboard cli into separate start / stop subcommands (#6234) ## Changes Made Follow up from https://github.com/Eventual-Inc/Daft/pull/5993 to split the cli args into subcommands to make it a bit neater to use. | 7 个月前 | |
chore: use dedicated OSS AWS account (#6442) ## Changes Made Update to use new dedicated OSS AWS account: - S3 buckets for GitHub artifacts, and public datasets - CloudFront distribution Requires: - Update ACTIONS_AWS_IAM_ROLE GitHub secret ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 6 个月前 | |
chore: use dedicated OSS AWS account (#6442) ## Changes Made Update to use new dedicated OSS AWS account: - S3 buckets for GitHub artifacts, and public datasets - CloudFront distribution Requires: - Update ACTIONS_AWS_IAM_ROLE GitHub secret ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> | 6 个月前 | |
[CHORE] Add TPC-H questions 11-22 to benchmarks (currently skipped) (#2299) | 2 年前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 7 个月前 | ||
| 6 个月前 | ||
| 6 个月前 | ||
| 6 个月前 | ||
| 6 个月前 | ||
| 6 个月前 | ||
| 7 个月前 | ||
| 6 个月前 | ||
| 3 年前 | ||
| 2 年前 | ||
| 7 个月前 | ||
| 6 个月前 | ||
| 6 个月前 | ||
| 2 年前 |