| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
fix(processor): allow crypto_stream ranges ending at memory limit (#3636) Co-authored-by: François Garillot <4142+huitseeker@users.noreply.github.com> | 1 个月前 | |
Reduce processor compilation time (#3292) * Relax source-aware execution inlining Relax the remaining source-aware processor node handlers from #[inline(always)] to #[inline]. This keeps the review-facing intent explicit while allowing LLVM to choose whether each generic helper should actually be inlined. I tried the more aggressive shape first. Broadly deleting these annotations and broadly replacing them with #[inline] gave the same benchmark-build proxy result, 1m50s versus 1m49s, so deletion was not needed for the compile-time effect. The narrow version keeps the pure and replay helpers at #[inline(always)] because the broad patch slowed trace building. Blake3 1-to-1 paired benchmark proof, using 1000 samples with a 3s measurement and 1s warmup: execute_trace_inputs_sync improved from 8.558320218 ms to 8.399991024 ms, -1.85%, comparing /tmp/miden-vm-issue-3271-exec-base/result.json with /tmp/miden-vm-issue-3271-exec-inline-narrow/result.json. build_trace improved from 56.800079456666666 ms to 55.13597942666667 ms, -2.93%, comparing /tmp/miden-vm-issue-3271-buildtrace-base/result.json with /tmp/miden-vm-issue-3271-buildtrace-inline-narrow/result.json. Compile proof: adjusted test-dev core-test --no-run build improved from 6m04s, real 364.72s, user 1276.43s, sys 38.89s, to 3m43s, real 223.90s, user 1009.24s, sys 34.10s. The no-override measurement improved from 24m34s, real 1475.13s, user 4461.06s, sys 65.76s, to 17m58s, real 1078.76s, user 3290.54s, sys 47.66s. The no-override build is still about 4.8x slower than the adjusted build, so this does not justify removing the profile override. Validation: make check; make test-fast; make lint; make check-features. * Relax basic block orchestration inlining Convert the basic block entry, resume, batch, and finish helpers from #[inline(always)] to #[inline]. This keeps an explicit inline hint while letting LLVM avoid forcing these larger generic orchestration functions into every caller. The per-operation execute_op_batch helper stays #[inline(always)], as do the pure/replay hot paths changed in earlier experiments. This keeps the change focused on orchestration functions rather than the inner operation loop. Compile proof, using fresh CARGO_TARGET_DIR builds of make core-test FEATURES="concurrent,testing" EXPR="-E 'not test(#*proptest) and not test(cli_)'" EXTRA=--no-run: baseline HEAD 04051d15ce took 3m40s, real 221.11s, user 993.00s, sys 32.88s; this patch took 1m17s, real 78.26s, user 523.18s, sys 32.87s. Blake3 1-to-1 paired light-axis proof, with 1000 samples, 3s measurement, 1s warmup, and axes execute_trace_inputs_sync,build_trace: execute_trace_inputs_sync improved from 8.397938371 ms to 8.084343752 ms; build_trace improved from 54.380447955 ms to 54.134897112 ms. Results are in /tmp/miden-vm-issue-3271-blake3-basic-block-baseline/result.json and /tmp/miden-vm-issue-3271-blake3-basic-block-candidate/result.json. Validation: make check; make test-fast; make lint; make check-features. * chore: Changelog | 2 个月前 | |
Reduce processor compilation time (#3292) * Relax source-aware execution inlining Relax the remaining source-aware processor node handlers from #[inline(always)] to #[inline]. This keeps the review-facing intent explicit while allowing LLVM to choose whether each generic helper should actually be inlined. I tried the more aggressive shape first. Broadly deleting these annotations and broadly replacing them with #[inline] gave the same benchmark-build proxy result, 1m50s versus 1m49s, so deletion was not needed for the compile-time effect. The narrow version keeps the pure and replay helpers at #[inline(always)] because the broad patch slowed trace building. Blake3 1-to-1 paired benchmark proof, using 1000 samples with a 3s measurement and 1s warmup: execute_trace_inputs_sync improved from 8.558320218 ms to 8.399991024 ms, -1.85%, comparing /tmp/miden-vm-issue-3271-exec-base/result.json with /tmp/miden-vm-issue-3271-exec-inline-narrow/result.json. build_trace improved from 56.800079456666666 ms to 55.13597942666667 ms, -2.93%, comparing /tmp/miden-vm-issue-3271-buildtrace-base/result.json with /tmp/miden-vm-issue-3271-buildtrace-inline-narrow/result.json. Compile proof: adjusted test-dev core-test --no-run build improved from 6m04s, real 364.72s, user 1276.43s, sys 38.89s, to 3m43s, real 223.90s, user 1009.24s, sys 34.10s. The no-override measurement improved from 24m34s, real 1475.13s, user 4461.06s, sys 65.76s, to 17m58s, real 1078.76s, user 3290.54s, sys 47.66s. The no-override build is still about 4.8x slower than the adjusted build, so this does not justify removing the profile override. Validation: make check; make test-fast; make lint; make check-features. * Relax basic block orchestration inlining Convert the basic block entry, resume, batch, and finish helpers from #[inline(always)] to #[inline]. This keeps an explicit inline hint while letting LLVM avoid forcing these larger generic orchestration functions into every caller. The per-operation execute_op_batch helper stays #[inline(always)], as do the pure/replay hot paths changed in earlier experiments. This keeps the change focused on orchestration functions rather than the inner operation loop. Compile proof, using fresh CARGO_TARGET_DIR builds of make core-test FEATURES="concurrent,testing" EXPR="-E 'not test(#*proptest) and not test(cli_)'" EXTRA=--no-run: baseline HEAD 04051d15ce took 3m40s, real 221.11s, user 993.00s, sys 32.88s; this patch took 1m17s, real 78.26s, user 523.18s, sys 32.87s. Blake3 1-to-1 paired light-axis proof, with 1000 samples, 3s measurement, 1s warmup, and axes execute_trace_inputs_sync,build_trace: execute_trace_inputs_sync improved from 8.397938371 ms to 8.084343752 ms; build_trace improved from 54.380447955 ms to 54.134897112 ms. Results are in /tmp/miden-vm-issue-3271-blake3-basic-block-baseline/result.json and /tmp/miden-vm-issue-3271-blake3-basic-block-candidate/result.json. Validation: make check; make test-fast; make lint; make check-features. * chore: Changelog | 2 个月前 | |
feat(debug): add initial support for debugging inline calls (#3427) * fix(debug): preserve source context while stepping * fix: add changelog input * feat(debug): emit inline calls with consolidated metadata * fix(debug): retain inline chains on MAST control nodes * fix(debug): preserve inline contexts across MAST uses * fix(debug): bound inline exec source growth * fix(debug): stabilize external inline contexts * fix(debug): preserve exact inline source contexts * fix(debug): harden inline source graph handling * test(assembly): reuse named inline marker helper * fix(assembly-syntax): propagate debug types serde | 1 个月前 | |
feat(debug): add initial support for debugging inline calls (#3427) * fix(debug): preserve source context while stepping * fix: add changelog input * feat(debug): emit inline calls with consolidated metadata * fix(debug): retain inline chains on MAST control nodes * fix(debug): preserve inline contexts across MAST uses * fix(debug): bound inline exec source growth * fix(debug): stabilize external inline contexts * fix(debug): preserve exact inline source contexts * fix(debug): harden inline source graph handling * test(assembly): reuse named inline marker helper * fix(assembly-syntax): propagate debug types serde | 1 个月前 | |
Reduce processor compilation time (#3292) * Relax source-aware execution inlining Relax the remaining source-aware processor node handlers from #[inline(always)] to #[inline]. This keeps the review-facing intent explicit while allowing LLVM to choose whether each generic helper should actually be inlined. I tried the more aggressive shape first. Broadly deleting these annotations and broadly replacing them with #[inline] gave the same benchmark-build proxy result, 1m50s versus 1m49s, so deletion was not needed for the compile-time effect. The narrow version keeps the pure and replay helpers at #[inline(always)] because the broad patch slowed trace building. Blake3 1-to-1 paired benchmark proof, using 1000 samples with a 3s measurement and 1s warmup: execute_trace_inputs_sync improved from 8.558320218 ms to 8.399991024 ms, -1.85%, comparing /tmp/miden-vm-issue-3271-exec-base/result.json with /tmp/miden-vm-issue-3271-exec-inline-narrow/result.json. build_trace improved from 56.800079456666666 ms to 55.13597942666667 ms, -2.93%, comparing /tmp/miden-vm-issue-3271-buildtrace-base/result.json with /tmp/miden-vm-issue-3271-buildtrace-inline-narrow/result.json. Compile proof: adjusted test-dev core-test --no-run build improved from 6m04s, real 364.72s, user 1276.43s, sys 38.89s, to 3m43s, real 223.90s, user 1009.24s, sys 34.10s. The no-override measurement improved from 24m34s, real 1475.13s, user 4461.06s, sys 65.76s, to 17m58s, real 1078.76s, user 3290.54s, sys 47.66s. The no-override build is still about 4.8x slower than the adjusted build, so this does not justify removing the profile override. Validation: make check; make test-fast; make lint; make check-features. * Relax basic block orchestration inlining Convert the basic block entry, resume, batch, and finish helpers from #[inline(always)] to #[inline]. This keeps an explicit inline hint while letting LLVM avoid forcing these larger generic orchestration functions into every caller. The per-operation execute_op_batch helper stays #[inline(always)], as do the pure/replay hot paths changed in earlier experiments. This keeps the change focused on orchestration functions rather than the inner operation loop. Compile proof, using fresh CARGO_TARGET_DIR builds of make core-test FEATURES="concurrent,testing" EXPR="-E 'not test(#*proptest) and not test(cli_)'" EXTRA=--no-run: baseline HEAD 04051d15ce took 3m40s, real 221.11s, user 993.00s, sys 32.88s; this patch took 1m17s, real 78.26s, user 523.18s, sys 32.87s. Blake3 1-to-1 paired light-axis proof, with 1000 samples, 3s measurement, 1s warmup, and axes execute_trace_inputs_sync,build_trace: execute_trace_inputs_sync improved from 8.397938371 ms to 8.084343752 ms; build_trace improved from 54.380447955 ms to 54.134897112 ms. Results are in /tmp/miden-vm-issue-3271-blake3-basic-block-baseline/result.json and /tmp/miden-vm-issue-3271-blake3-basic-block-candidate/result.json. Validation: make check; make test-fast; make lint; make check-features. * chore: Changelog | 2 个月前 | |
Reduce processor compilation time (#3292) * Relax source-aware execution inlining Relax the remaining source-aware processor node handlers from #[inline(always)] to #[inline]. This keeps the review-facing intent explicit while allowing LLVM to choose whether each generic helper should actually be inlined. I tried the more aggressive shape first. Broadly deleting these annotations and broadly replacing them with #[inline] gave the same benchmark-build proxy result, 1m50s versus 1m49s, so deletion was not needed for the compile-time effect. The narrow version keeps the pure and replay helpers at #[inline(always)] because the broad patch slowed trace building. Blake3 1-to-1 paired benchmark proof, using 1000 samples with a 3s measurement and 1s warmup: execute_trace_inputs_sync improved from 8.558320218 ms to 8.399991024 ms, -1.85%, comparing /tmp/miden-vm-issue-3271-exec-base/result.json with /tmp/miden-vm-issue-3271-exec-inline-narrow/result.json. build_trace improved from 56.800079456666666 ms to 55.13597942666667 ms, -2.93%, comparing /tmp/miden-vm-issue-3271-buildtrace-base/result.json with /tmp/miden-vm-issue-3271-buildtrace-inline-narrow/result.json. Compile proof: adjusted test-dev core-test --no-run build improved from 6m04s, real 364.72s, user 1276.43s, sys 38.89s, to 3m43s, real 223.90s, user 1009.24s, sys 34.10s. The no-override measurement improved from 24m34s, real 1475.13s, user 4461.06s, sys 65.76s, to 17m58s, real 1078.76s, user 3290.54s, sys 47.66s. The no-override build is still about 4.8x slower than the adjusted build, so this does not justify removing the profile override. Validation: make check; make test-fast; make lint; make check-features. * Relax basic block orchestration inlining Convert the basic block entry, resume, batch, and finish helpers from #[inline(always)] to #[inline]. This keeps an explicit inline hint while letting LLVM avoid forcing these larger generic orchestration functions into every caller. The per-operation execute_op_batch helper stays #[inline(always)], as do the pure/replay hot paths changed in earlier experiments. This keeps the change focused on orchestration functions rather than the inner operation loop. Compile proof, using fresh CARGO_TARGET_DIR builds of make core-test FEATURES="concurrent,testing" EXPR="-E 'not test(#*proptest) and not test(cli_)'" EXTRA=--no-run: baseline HEAD 04051d15ce took 3m40s, real 221.11s, user 993.00s, sys 32.88s; this patch took 1m17s, real 78.26s, user 523.18s, sys 32.87s. Blake3 1-to-1 paired light-axis proof, with 1000 samples, 3s measurement, 1s warmup, and axes execute_trace_inputs_sync,build_trace: execute_trace_inputs_sync improved from 8.397938371 ms to 8.084343752 ms; build_trace improved from 54.380447955 ms to 54.134897112 ms. Results are in /tmp/miden-vm-issue-3271-blake3-basic-block-baseline/result.json and /tmp/miden-vm-issue-3271-blake3-basic-block-candidate/result.json. Validation: make check; make test-fast; make lint; make check-features. * chore: Changelog | 2 个月前 | |
feat(debug): add initial support for debugging inline calls (#3427) * fix(debug): preserve source context while stepping * fix: add changelog input * feat(debug): emit inline calls with consolidated metadata * fix(debug): retain inline chains on MAST control nodes * fix(debug): preserve inline contexts across MAST uses * fix(debug): bound inline exec source growth * fix(debug): stabilize external inline contexts * fix(debug): preserve exact inline source contexts * fix(debug): harden inline source graph handling * test(assembly): reuse named inline marker helper * fix(assembly-syntax): propagate debug types serde | 1 个月前 | |
Reduce processor compilation time (#3292) * Relax source-aware execution inlining Relax the remaining source-aware processor node handlers from #[inline(always)] to #[inline]. This keeps the review-facing intent explicit while allowing LLVM to choose whether each generic helper should actually be inlined. I tried the more aggressive shape first. Broadly deleting these annotations and broadly replacing them with #[inline] gave the same benchmark-build proxy result, 1m50s versus 1m49s, so deletion was not needed for the compile-time effect. The narrow version keeps the pure and replay helpers at #[inline(always)] because the broad patch slowed trace building. Blake3 1-to-1 paired benchmark proof, using 1000 samples with a 3s measurement and 1s warmup: execute_trace_inputs_sync improved from 8.558320218 ms to 8.399991024 ms, -1.85%, comparing /tmp/miden-vm-issue-3271-exec-base/result.json with /tmp/miden-vm-issue-3271-exec-inline-narrow/result.json. build_trace improved from 56.800079456666666 ms to 55.13597942666667 ms, -2.93%, comparing /tmp/miden-vm-issue-3271-buildtrace-base/result.json with /tmp/miden-vm-issue-3271-buildtrace-inline-narrow/result.json. Compile proof: adjusted test-dev core-test --no-run build improved from 6m04s, real 364.72s, user 1276.43s, sys 38.89s, to 3m43s, real 223.90s, user 1009.24s, sys 34.10s. The no-override measurement improved from 24m34s, real 1475.13s, user 4461.06s, sys 65.76s, to 17m58s, real 1078.76s, user 3290.54s, sys 47.66s. The no-override build is still about 4.8x slower than the adjusted build, so this does not justify removing the profile override. Validation: make check; make test-fast; make lint; make check-features. * Relax basic block orchestration inlining Convert the basic block entry, resume, batch, and finish helpers from #[inline(always)] to #[inline]. This keeps an explicit inline hint while letting LLVM avoid forcing these larger generic orchestration functions into every caller. The per-operation execute_op_batch helper stays #[inline(always)], as do the pure/replay hot paths changed in earlier experiments. This keeps the change focused on orchestration functions rather than the inner operation loop. Compile proof, using fresh CARGO_TARGET_DIR builds of make core-test FEATURES="concurrent,testing" EXPR="-E 'not test(#*proptest) and not test(cli_)'" EXTRA=--no-run: baseline HEAD 04051d15ce took 3m40s, real 221.11s, user 993.00s, sys 32.88s; this patch took 1m17s, real 78.26s, user 523.18s, sys 32.87s. Blake3 1-to-1 paired light-axis proof, with 1000 samples, 3s measurement, 1s warmup, and axes execute_trace_inputs_sync,build_trace: execute_trace_inputs_sync improved from 8.397938371 ms to 8.084343752 ms; build_trace improved from 54.380447955 ms to 54.134897112 ms. Results are in /tmp/miden-vm-issue-3271-blake3-basic-block-baseline/result.json and /tmp/miden-vm-issue-3271-blake3-basic-block-candidate/result.json. Validation: make check; make test-fast; make lint; make check-features. * chore: Changelog | 2 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 1 个月前 | ||
| 2 个月前 | ||
| 2 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 2 个月前 | ||
| 2 个月前 | ||
| 1 个月前 | ||
| 2 个月前 |