已开启
【WIP】LRU使能shard模式+新增meta_service_benchmark #401
chenxin创建于 7月23日
【WIP】LRU使能shard模式+新增meta_service_benchmark #401
已开启
chenxin创建于 7月23日
chenxin
chenxin成员
7月23日

合入来源

问题/功能描述

进一步可以考虑benchmark中,get只读50%占比的热数据,put和evict操作50%的冷数据,可能可以获得更大收益
还可以手撕哈希表+双向链表,提升访问性能

修改方案描述

image.png
8个client,每个client起4个线程,alloc时延在250~350us左右
1个client,只起1个线程,alloc时延在22us左右

image.png
拆分为16个shard后,alloc耗时减少一半多,PUT性能提升一倍多

是否涉及UT/ST

开发自检

likedislike
合并受阻
ascend-robot
ascend-robot成员
7月23日 评论:
流水线 PR-pipeline_memcache#1675 [ commitID:82302238 ] 运行失败
阶段 任务名 状态 详情
编译构建 Build_memcache 🕚 >>>
恶意代码检查 Antipoison_memcache 🕚 >>>
编码安全与规范检查 pre-commit ❌ >>>
CodeCheck_memcache 🕚 >>>
开源片段检查 SCA_memcache 🕚 >>>
开发者测试 UT_memcache 🕚 >>>
流水线 PR-pipeline_memcache ❌ >>>
此流水线已支持下列评论快捷指令,仅PR创建者和白名单成员评论有效
  • compile : 运行流水线
  • retry : 重试流水线所有失败子任务
  • retry <任务名> : 仅重试指定失败子任务
  • stop : 停止流水线
likedislike
Xxiangjie10成员
7月23日 添加了label:pr-audit-failed
xiangjie10成员
7月23日 评论:
🔍 PR 规范审计未通过,以下项目需要修正:
  • ❌ PR 未关联里程碑或 Issue
  • ❌ PR 新增代码 1627 行超过 1000 行,且标题未标注"反合"

请修正后重新提交,或联系仓库管理员。

likedislike
ascend-robotascend-robot成员
7月23日 添加了label:stat/needs-squash
ascend-robotascend-robot成员
7月23日 添加了label:ascend-cla/yes
ascend-robot
ascend-robot成员
7月23日 评论:

Thanks for your pull-request.
The full list of commands accepted by me can be found at here。
You can get sig-info at here


PR Approval Progress

⚠️ This PR does not yet meet the following requirements:lgtm (requires ≥ 3 person(s) per module)、approve (requires ≥ 1 person(s) per module)

Module Approval Details

module lgtm status approve status
repo-Ascend/memcache ❌ (0/3)(You can also ask: wlwen, shilinlee_com, youngzhou66, zja4gitcode, p_chenhui) ❌ (0/1)(You can also ask: weihaoran1, 康富安, 程俊华, wlwen, shanyuelanhua)

💡 Tip:

  • Committer can comment /approve or /lgtm
  • Commenting /approve implies both code review (lgtm) and intent to merge (approve)

CLA Signature Pass

chenyz6, thanks for your pull request. All authors of the commits have signed the CLA. 👍

likedislike
chenxinchenxin成员
7月23日 修改标题为 “【WIP】LRU使能shard模式+新增meta_service_benchmark”,原标题为“【WIP】LRU使能shard模式”
此处折叠了11条事件消息 查看更多
chenxinchenxin成员
8月31日 修改了pull request 的描述
yyang728成员17 天前进行代码检视1
example/meta_service_bench/bench_meta_service.cpp
@@ -0,0 +154,4 @@
154+ int len = vsnprintf(buf, sizeof(buf), fmt, args);
155+ va_end(args);
156+ if (len > 0) {
157+ write(gTermFd, buf, static_cast<size_t>(len));
yyang72817 天前评论:

example/meta_service_bench/bench_meta_service.cpp:157

🔴 需验证: TermLog 中 write 的第三个参数 len 可能超过 buf 的大小,导致栈缓冲区越界读。

推导过程:

  • buf 类型为 char[1024],栈上分配。
  • vsnprintf(buf, sizeof(buf), fmt, args) 的返回值 len 是「若缓冲区足够大时本应写入的字符数」(不含 \0),而非实际写入的字符数。
  • 当格式化后的字符串长度 ≥ 1024 时,len ≥ 1024,但 buf 仅有 1024 字节。
  • 此时 write(gTermFd, buf, static_cast<size_t>(len)) 会从 buf 起始地址读取 len 字节,超出缓冲区边界,构成栈上的 buffer over-read。

触发路径: TermLog 在多处被调用(如行 454、498、508、554、590 等),若 fmt 参数拼接出 ≥ 1024 字符的消息即可触发。

建议修改:

if (len > 0) {
    size_t writeLen = std::min(static_cast<size_t>(len), sizeof(buf) - 1);
    write(gTermFd, buf, writeLen);
}
likedislike
yyang728成员17 天前进行代码检视1
example/meta_service_bench/bench_meta_service.cpp
@@ -0,0 +331,4 @@
331+ allocReq.operateId_ = operateId;
332+ AllocResponse allocResp;
333+ if (client.SyncCall(allocReq, allocResp, kRpcTimeoutMs) != MMC_OK) return MMC_ERROR;
334+ if (allocResp.result_ != MMC_OK || allocResp.blobs_.empty()) return allocResp.result_;
yyang72817 天前评论:

example/meta_service_bench/bench_meta_service.cpp:334

🟠 问题: 当 allocResp.result_ == MMC_OK 但 allocResp.blobs_ 为空时,此行返回 allocResp.result_(即 MMC_OK),导致调用方误认为 Put 成功,而实际未分配到任何 blob。

推导过程:

  • 条件 allocResp.result_ != MMC_OK || allocResp.blobs_.empty() 使用 || 连接。
  • 当 result_ == MMC_OK 且 blobs_.empty() == true 时,条件为 false || true = true,进入 return 分支。
  • 返回值为 allocResp.result_,即 MMC_OK。
  • 调用方 Worker 函数(行 425)判断 ret == MMC_OK 后将此操作计入 successCount,但实际并未完成数据写入。

影响: 在服务端返回成功但未分配 blob 的边界情况下,benchmark 统计的 Put 成功数会偏高,影响基准测试结果的准确性。

建议修改:

if (allocResp.result_ != MMC_OK) return allocResp.result_;
if (allocResp.blobs_.empty()) return MMC_ERROR;
likedislike