| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
Testing/architecture adapter cleanup (#1456) * initial sweep * Cleanup and improve tests to remove tautologies * restore gqa safety tests * Refactor and comment cleanup * Resolving CI failures | 3 个月前 | |
Fixing edge cases on heterogeneous configs (#1679) | 1 个月前 | |
Resolution for #112 and #830 (#1304) * Resolution for #644 and #341 * Started activation cache improvement * Full resolution for 210 + a demo notebook * Resolution for #796, Factored Matrix memory leak * Resolved #453, underlying issue * Resolution for Issue #385, added notes about forced eager, added a test to check for future drift * Added hook introspection mixin for #297 * Made improvements to booting training revisions * Adapter test improvements * format cleanup * Adding a way to display logit vector for #112 and add type hinting for #830 * Type issue resolution – Resolving import confusion between TransformerBridgeConfig module and class, which were separate entities sharing the same name. Files renamed to properly differentiate * Format and type fixes * Removed unnecssary assertion | 4 个月前 | |
Chore/comment cleanup (#1674) * Hook extensions for specialized bridges * Additional bug fixes * Comment cleanup | 1 个月前 | |
fix(neox): resolve unembedding as lm_head for transformers >= 5.13 (#1752) * fix(neox): resolve unembedding as lm_head for transformers >= 5.13 - Map the NeoX bridge unembed component to `lm_head` instead of the removed `embed_out`, so `TransformerBridge.boot_transformers` boots GPT-NeoX/Pythia on the pinned transformers 5.13.0 instead of raising `AttributeError: 'GPTNeoXForCausalLM' object has no attribute 'embed_out'`. - Mirror the same rename on the HookedTransformer weight-conversion path so the bridge and legacy systems stay consistent (AGENTS §2). - Found while generalizing Backward Lens to dense-MLP families, which needs pythia-70m as an `out_in` integration model. * test(neox): cover unembedding resolution on the transformers >= 5.13 layout - Update the top-level HF-path assertion to expect the `lm_head` unembed name. - Add a model-free regression test that resolves the unembed component against a module exposing `lm_head` (and not `embed_out`) via `get_remote_component`, guarding against reintroducing the stale name that broke GPT-NeoX/Pythia boot. * fix(neox): resolve unembedding on both embed_out and lm_head layouts CI's locked transformers==5.13.0 still exposes embed_out; the earlier lm_head-only rename only matched later transformers versions (5.15.1 in the local conda env), breaking every pythia/GPT-NeoX boot in CI. Resolve the unembed name at runtime (bridge: ArchitectureAdapter.prepare_model once the real HF module is available; HookedTransformer: hasattr in convert_neox_weights) so both layouts work. * fix(neox): guard component_mapping before unembed lookup - assert component_mapping is not None in prepare_model before indexing "unembed" - narrows ComponentMapping | None so mypy allows the subscript (fixes index error) * fix(neox): correct transformers embed_out→lm_head rename boundary to 5.14.0 - The GPTNeoXForCausalLM embed_out→lm_head rename landed in transformers 5.14.0 (huggingface/transformers#47198), not 5.13; fix the docstring/comment/test references that said 5.13 or ~5.14 - Rename test_unembed_resolves_against_transformers_5_13_lm_head_layout to _5_14_ so it no longer contradicts the locked-5.13.0-still-has-embed_out fallback test - Leave the four "locked 5.13.0 still exposes embed_out" notes unchanged — that pairing is correct * refactor(neox): set unembed name via components accessor in prepare_model - Replace the redundant `assert component_mapping is not None` + `isinstance(unembed, UnembeddingBridge)` guard with a direct `self.components["unembed"].name = "embed_out"` assignment. - The `components` accessor already asserts the mapping is built and `.name` is declared on GeneralizedComponent, matching how bert.py and other adapters edit their mapping; keep the hasattr `if` guard as-is. | 26 天前 | |
Fix gpt oss olmo3 parity (#1621) * Resolution for issue 1619 * Updated for 1620 * Add verification script * Fixed 1619 on HF * Verification script repair * fixing per-layer olmo * cleanup | 1 个月前 | |
Chore/comment cleanup (#1674) * Hook extensions for specialized bridges * Additional bug fixes * Comment cleanup | 1 个月前 | |
Refuse quantized weights in weight-space code paths instead of reading them as matrices (#1669) * initial numerics fix * Additional clarification and bug cleanup * Fixing issues with MoE and Dense hooks * test cleanup and improvements * Mixtral 5.x bug fix * Quantization bugs * Missing tests restored * Cleanup of potential errors * fix formatting | 1 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 3 个月前 | ||
| 1 个月前 | ||
| 4 个月前 | ||
| 1 个月前 | ||
| 26 天前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 1 个月前 |