| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
chore: reduce binary size by feature flagging derive(Debug) (#5622) ## Changes Made cargo-llvm-lines showed that 0.7% of our binary size was from Debug impls 40559 (0.7%, 5.8%) 15604 (3.0%, 5.4%) <&T as core::fmt::Debug>::fmt I noticed that this PR reduces the total binary size by about 1.2MB. While pretty small when our uncompressed binary is 180+MB, the small savings add up. So this PR attempts to either remove derive(Debug) or feature flag it: cfg_attr(debug_assertions, derive(Debug)) in some spots we still need it to not break a bunch of things, so for not(debug_assertions), theres a simpler debug impl provided, that usually just debugs the name, not all of the struct fields. ## Related Issues <!-- Link to related GitHub issues, e.g., "Closes #123" --> ## Checklist - [ ] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [ ] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [ ] Documentation builds and is formatted properly | 10 个月前 | |
feat: implements an openai provider with embed_text (#4997) ## Changes Made * Adds an OpenAI provider implementation along with provider-specific options. * Adds an OpenAI TextEmbedder with dynamic batching support. * Adds additional session methods for the providers e.g. set_provider, get_provider * Adds support for attaching custom providers with proper resolution. * Makes set_provider initialize with defaults e.g. daft.set_provider("openai") * Adds a with daft.session() context manager for scoped access to sessions. * Documents new APIs 😄 ## Examples python import daft with daft.session() as sess: sess.set_provider("openai") # everything within this block resolves via the scoped session. python In [1]: import daft In [2]: from daft.functions.ai import embed_text In [3]: daft.set_provider("openai") In [5]: df = daft.from_pydict({"text": ["hello, world!"]}) In [6]: df = df.with_column("embedding", embed_text(df["text"])) In [7]: df.show() ╭───────────────┬──────────────────────────╮ │ text ┆ embedding │ │ --- ┆ --- │ │ Utf8 ┆ Embedding[Float32; 1536] │ ╞═══════════════╪══════════════════════════╡ │ hello, world! ┆ <Embedding> │ ╰───────────────┴──────────────────────────╯ (Showing first 1 of 1 rows) Something interesting is that daft's context is a global singleton, and it's not clear if we would ever want to enable different. If yes, then this PR is a step towards that. If not, then I'll switch the context manager to be for the session only. We have talked in the past about using a context manager to set the error-handling behavior, which is why I've gone the daft.use_context angle. ## TODO * [x] Fix mypy * [x] Unit test * [x] Improve OpenAI embed_text implementation ## Checklist - [x] Documented in API Docs (if applicable) - [ ] Documented in User Guide (if applicable) - [x] If adding a new documentation page, doc is added to docs/mkdocs.yml navigation - [x] Documentation builds and is formatted properly (tag @/ccmao1130 for docs review) | 1 年前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 10 个月前 | ||
| 1 年前 |