| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
docs: Align MultiVectorEncoder pages with the other archetypes (#3933) * docs: Align MultiVectorEncoder pages with the other archetypes The MultiVectorEncoder documentation had drifted from the shape of the Sentence Transformer, Cross Encoder, and Sparse Encoder pages. usage.rst gains the modality section its siblings have: the modalities / supports lead-in, the extras tip, the accepted input formats, and a Modality Support sidebar. Code blocks are normalized to a single indentation and directive style, the query / document methods get a worked example, and the page is reordered into the sibling shape. training_overview.md gains a Comparisons with SentenceTransformer Training section and a Multimodal Datasets subsection, and the Model section is split into tabs like its siblings. The visual document retrieval tab now leads with loading a multimodal (or omnimodal) embedding backbone, with the transformers-native *ForRetrieval checkpoints kept as the compatibility route rather than the headline. pretrained_models.md is rebuilt from the blogpost tables: 25 text and 20 visual checkpoints with parameter counts, backbones, and the revision / trust_remote_code notes each one needs. custom_models.rst gets Title Case headings and a Loading section to pair with Saving, and a new examples/multi_vector_encoder/evaluation/README.md makes the evaluation script reachable from the docs tree. Examples now showcase LateOn and mLateOn, with LFM2-ColBERT-350M, mxbai-edge-colbert-v0-32m, answerai-colbert-small-v1, and pplx-embed-v1-late-0.6b covering the other load paths. Every printed output in the docs was measured rather than copied. * similar as -> similar to Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> * docs: Quiet the Notes column, flag refs/pr/N for release --------- Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> | 1 个月前 | |
[docs] Name the metric on the quality axis of the backend benchmark figures (#3977) The bottom subfigure was labelled "Performance Ratio", which reads as a second speed measure rather than the retrieval quality it plots. It now names the metrics that actually fed it: NanoBEIR NDCG@10 for the Cross Encoder, Sparse Encoder and Multi-Vector Encoder figures, and NanoBEIR NDCG@10 plus STSb Spearman for the Sentence Transformer ones. The figure title and the surrounding prose say "quality" instead of "performance" to match. All eight figures are re-rendered from unchanged benchmark results, so every number is identical to before. Closes #3950 | 27 天前 | |
docs: Document shared input formats across model archetypes (#4036) * Document shared input formats across model archetypes * Clarify input processing guidance and audio resampling | 12 天前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 1 个月前 | ||
| 27 天前 | ||
| 12 天前 |