| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
refactor(inference): unify request handling across entrypoints | 25 天前 | |
update doc | 1 年前 | |
update doc | 1 年前 | |
update doc | 1 年前 | |
Update offload.md (#1069) | 4 个月前 | |
update parallel doc | 1 年前 | |
refactor(api)!: unify inference parameters and remove RIFE (#1511) | 24 天前 | |
refactor: remove bundled Gradio UI and isolate SekoTalk resize handling (#1532) | 16 天前 | |
add support of torch profiling (#1101) ### **Summary** Add a reusable PyTorch Profiler trace utility for LightX2V and document how to use it. Profiling is opt-in at the call site via TorchTraceProfileContext; normal inference is unchanged when profiling is not enabled. This MR replaces the ad-hoc, Qwen-specific profiling path in transformer_infer.py with a centralized helper that supports TensorBoard and Chrome trace export. ### **Changes** **New module** lightx2v/utils/torch_trace_profiler.py - TorchTraceProfileContext: wrap any callable at its call site - Configurable schedule (wait / warmup / active), export format, output paths, and optional Python stack traces - Exports to TensorBoard (.pt.trace.json) or Chrome trace (.json) - One profile session per process; first invoked call site wins, others run normally with a warning **Documentation** docs/ZH_CN/source/method_tutorials/torch_profiling.md - Usage guide, parameter reference, TensorBoard / Perfetto viewing instructions, and SSH port-forwarding notes - Linked from docs/ZH_CN/source/index.rst **Qwen Image integration** - Remove _infer_calculating_profiled() from qwen_image/infer/transformer_infer.py - Add a commented usage example in qwen_image_runner.py around infer_main (uncomment import + context to enable) ### **Motivation** Kernel-level profiling is useful for Magi compile tuning and general performance work, but the previous approach was embedded inside the transformer and tied to a fixed Chrome trace export. The new utility keeps profiling at the runner/call-site layer, works across models, and supports both TensorBoard and Chrome formats without affecting production inference paths. ### **How to use** 1. Uncomment in qwen_image_runner.py: `from lightx2v.utils.torch_trace_profiler import TorchTraceProfileContext` 2. Wrap the target call, e.g.: ``` with TorchTraceProfileContext("🚀 infer_main", tb_dir="save_results/torch_profile", with_stack=True) as profile: profile.run(self.model.infer, self.inputs) ``` 3. Run inference as usual; check logs for trace paths and viewing commands. See the tutorial doc for full details. ### **Test plan** - Run Qwen Image I2I without profiling enabled — behavior and output unchanged - Enable TorchTraceProfileContext on infer_main — trace files written to save_results/torch_profile/ - Open TensorBoard → PYTORCH PROFILER tab and verify CPU/CUDA kernels are visible - Repeat with profile_format="chrome" — export loads in Perfetto - Confirm a second profile call site in the same process is skipped with a warning - Verify distributed runs (if applicable): only the first rank/profiled call site exports as expected ### **Notes** - No new config flags or env vars; profiling is controlled entirely by uncommenting/wrapping at the call site - Default output: {cwd}/save_results/torch_profile (TensorBoard) or {cwd}/save_results/trace.json (Chrome) - Requires tensorboard and torch-tb-profiler for TensorBoard viewing (documented in the tutorial) --------- Co-authored-by: Cursor <cursoragent@cursor.com> | 4 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 25 天前 | ||
| 1 年前 | ||
| 1 年前 | ||
| 1 年前 | ||
| 4 个月前 | ||
| 1 年前 | ||
| 24 天前 | ||
| 16 天前 | ||
| 4 个月前 |