aclnnFfnWorkerBatching
产品支持情况
- Ascend 950PR/Ascend 950DT:支持
- Atlas A3 训练系列产品/Atlas A3 推理系列产品:支持
- Atlas A2 训练系列产品/Atlas A2 推理系列产品:支持
- Atlas 200I/500 A2 推理产品:不支持
- Atlas 推理系列产品:不支持
- Atlas 训练系列产品:不支持
功能说明
-
接口功能:Attention与FFN分离部署场景下,FFN worker侧的token重排算子。Attention将token按专家路由发送到对应FFN worker的预分配数据区,本接口从该数据区中扫描调度信息,按专家维度聚合并重排token,产出各专家对应的连续token数据块。
该算子不建议单独使用,建议与FfnWorkerScheduler等算子配合使用,形成完整的工作流。
-
接收Attention侧发送的数据。该数据以ScheduleContext结构体内存排布方式存储。其具体定义参见调用示例。该结构体包含CommonArea、ControlArea、AttentionArea、FfnArea域。本接口从FfnArea中读取token_info_buf和token_data_buf获取待重排的token数据与描述信息,并获取layer_id、session_id、micro_batch_id、expert_ids等路由信息。
-
对
expert_ids_in中所有token的专家ID进行排序(被mask的token初始化为大值),生成gather索引。 -
多核并行按gather索引从
token_data中提取token的hidden states和dynamic scale,同时查表得到对应的session_id、micro_batch_id、token_id。 -
单核扫描排序后的专家ID序列,查找跳变点,生成
group_list(每个专家处理的token起止偏移)。
其中 Y=A×BS×(K+1)Y = A \times BS \times (K+1),AA 为Attention worker数量,BSBS 为micro batch size,K+1K+1 为topK加共享专家数。
-
-
计算公式:
gather_idx,_=sort(expert_ids_in)\text{gather\_idx}, \_ = \text{sort}(\text{expert\_ids\_in})
y[i]=token_data[gather_idx[i]]y[i] = \text{token\_data}[\text{gather\_idx}[i]]
group_list[e]=[expert_ide,expert_token_nume]\text{group\_list}[e] = [\text{expert\_id}_e, \text{expert\_token\_num}_e]
actual_token_num=∑eexpert_token_nume\text{actual\_token\_num} = \sum_{e} \text{expert\_token\_num}_e
函数原型
每个算子分为两段式接口,必须先调用“aclnnFfnWorkerBatchingGetWorkspaceSize”接口获取计算所需workspace大小以及包含了算子计算流程的执行器,再调用“aclnnFfnWorkerBatching”接口执行计算。
aclnnStatus aclnnFfnWorkerBatchingGetWorkspaceSize(
const aclTensor *scheduleContext,
int64_t expertNum,
const aclIntArray *maxOutShape,
int64_t tokenDtype,
int64_t needSchedule,
int64_t layerNum,
const aclTensor *y,
const aclTensor *groupList,
const aclTensor *sessionIds,
const aclTensor *microBatchIds,
const aclTensor *tokenIds,
const aclTensor *expertOffsets,
const aclTensor *dynamicScale,
const aclTensor *actualTokenNum,
uint64_t *workspaceSize,
aclOpExecutor **executor)
aclnnStatus aclnnFfnWorkerBatching(
void* workspace,
uint64_t workspaceSize,
aclOpExecutor* executor,
const aclrtStream stream)
aclnnFfnWorkerBatchingGetWorkspaceSize
-
参数说明
参数名 输入/输出 描述 使用说明 数据类型 数据格式 维度(shape) 非连续Tensor scheduleContext 输入 FFN侧接收的调度上下文,内含CommonArea、ControlArea、AttentionArea、FfnArea。算子从FfnArea中读取token_info_buf和token_data_buf获取待重排的token数据与描述信息。详细结构参见调用示例。 不支持空tensor。 INT8 ND 1维,shape固定为(1024) × expertNum 输入 本卡专家总数,等于每层本卡专家数 × layer_num。用于推导group_list输出大小。 取值范围为(0, 8192]。 INT64 - - - maxOutShape 输入 输出shape上限,格式为 {A, BS, topK+1, H}。用于推导y输出的shape上限 Y = A × BS × (topK+1),以及H值。 数组长度必须为4。其中A取值范围为(0, 1024],BS大于0,topK+1取值范围为(0, 64],H大于0。 LIST_INT64 - - - tokenDtype 输入 输入token的数据类型。0表示FP16;1表示BF16;2表示INT8动态量化(INT8数据与FP32 dynamic scale连续排布)。取值为2时需输出dynamic_scale。 取值为0、1或2。 INT64 - - - needSchedule 输入 调度模式。0表示仅做batching不扫描数据;1表示先扫描数据再做batching。 取值为0或1。 INT64 - - - layerNum 输入 层数,每层专家独立索引。 取值范围为[0, expertNum]。 INT64 - - - y 输出 重排后的token hidden states,按专家ID排序后连续存放。 - FP16、BF16、INT8 ND 2维,(Y, H),其中Y = A × BS × (topK+1) × groupList 输出 每个专家处理的token范围。每行格式为 [expert_id, expert_token_num],未使用的专家填 [0, 0]。 - INT64 ND 2维,(expertNum, 2) × sessionIds 输出 每个输出token对应的Attention session ID。 - INT32 ND 1维,(Y) × microBatchIds 输出 每个输出token对应的micro batch ID。 - INT32 ND 1维,(Y) × tokenIds 输出 每个输出token在原始输入中的位置索引。 - INT32 ND 1维,(Y) × expertOffsets 输出 每个输出token在其所属专家分组内的偏移。 - INT32 ND 1维,(Y) × dynamicScale 输出 动态量化的scale值,仅在tokenDtype=2时有效。tokenDtype为0或1时为空tensor。 - FP32 ND 1维,(Y) × actualTokenNum 输出 所有专家有效token数之和。 - INT64 ND 1维,(1) × workspaceSize 输出 返回需要在Device侧申请的workspace大小。 - - - - - executor 输出 返回op执行器,包含了算子计算流程。 - - - - - -
返回值:
aclnnStatus:返回状态码,具体参见aclnn返回码。
第一段接口会完成入参校验,出现以下场景时报错:
返回值 错误码 描述 ACLNN_ERR_PARAM_NULLPTR 161001 输入是空指针。 ACLNN_ERR_PARAM_INVALID 161002 输入数据类型不在支持的范围内。
aclnnFfnWorkerBatching
-
参数说明
参数名 输入/输出 描述 workspace 输入 在Device侧申请的workspace内存地址。 workspaceSize 输入 在Device侧申请的workspace大小,由第一段接口aclnnFfnWorkerBatchingGetWorkspaceSize获取。 executor 输入 op执行器,包含了算子计算流程。 stream 输入 指定执行任务的Stream。 -
返回值:
aclnnStatus:返回状态码,具体参见aclnn返回码。
约束说明
- 参数A(Attention worker数量)支持 ≤ 1024。
- 参数M(micro batch数量)支持 ≤ 64。
- 参数K(topK数)支持 ≤ 64。
- 参数BS(micro batch size)和Y支持泛化,无硬上限(受内存限制)。
- 参数H(hidden size)支持泛化。
- tokenDtype为2时,输入int8数据与fp32 scale连续排布。
- 确定性计算:
- aclnnFfnWorkerBatching默认确定性实现。
调用示例
示例代码如下,仅供参考,具体编译和执行过程请参考编译与运行样例。
#include <iostream>
#include <vector>
#include <memory>
#include <cstring>
#include "acl/acl.h"
#include "aclnnop/aclnn_ffn_worker_batching.h"
#define CHECK_RET(cond, return_expr) \
do { \
if (!(cond)) { \
return_expr; \
} \
} while (0)
#define LOG_PRINT(message, ...) \
do { \
printf(message, ##__VA_ARGS__); \
} while (0)
int64_t GetShapeSize(const std::vector<int64_t>& shape) {
int64_t shapeSize = 1;
for (auto i : shape) {
shapeSize *= i;
}
return shapeSize;
}
int Init(int32_t deviceId, aclrtStream* stream) {
auto ret = aclInit(nullptr);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclInit failed. ERROR: %d\n", ret); return ret);
ret = aclrtSetDevice(deviceId);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtSetDevice failed. ERROR: %d\n", ret); return ret);
ret = aclrtCreateStream(stream);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtCreateStream failed. ERROR: %d\n", ret); return ret);
return 0;
}
void Finalize(int32_t deviceId, aclrtStream stream) {
aclrtDestroyStream(stream);
aclrtResetDevice(deviceId);
aclFinalize();
}
template <typename T>
int CreateAclTensor(const std::vector<T>& hostData, const std::vector<int64_t>& shape, void** deviceAddr,
aclDataType dataType, aclTensor** tensor) {
auto size = GetShapeSize(shape) * sizeof(T);
auto ret = aclrtMalloc(deviceAddr, size, ACL_MEM_MALLOC_HUGE_FIRST);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtMalloc failed. ERROR: %d\n", ret); return ret);
ret = aclrtMemcpy(*deviceAddr, size, hostData.data(), size, ACL_MEMCPY_HOST_TO_DEVICE);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtMemcpy failed. ERROR: %d\n", ret); return ret);
std::vector<int64_t> stride(shape.size(), 1);
for (int64_t i = shape.size() - 2; i >= 0; i--) {
stride[i] = shape[i + 1] * stride[i + 1];
}
*tensor = aclCreateTensor(shape.data(), shape.size(), dataType, stride.data(), 0, aclFormat::ACL_FORMAT_ND,
shape.data(), shape.size(), *deviceAddr);
return 0;
}
int CreateAclTensorNoData(const std::vector<int64_t>& shape, void** deviceAddr, aclDataType dataType,
aclTensor** tensor) {
auto size = GetShapeSize(shape) * sizeof(int8_t);
if (dataType == ACL_INT32) { size = GetShapeSize(shape) * sizeof(int32_t); }
if (dataType == ACL_INT64) { size = GetShapeSize(shape) * sizeof(int64_t); }
if (dataType == ACL_FLOAT) { size = GetShapeSize(shape) * sizeof(float); }
if (dataType == ACL_FLOAT16) { size = GetShapeSize(shape) * sizeof(int16_t); }
auto ret = aclrtMalloc(deviceAddr, size, ACL_MEM_MALLOC_HUGE_FIRST);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtMalloc failed. ERROR: %d\n", ret); return ret);
std::vector<int64_t> stride(shape.size(), 1);
for (int64_t i = shape.size() - 2; i >= 0; i--) {
stride[i] = shape[i + 1] * stride[i + 1];
}
*tensor = aclCreateTensor(shape.data(), shape.size(), dataType, stride.data(), 0, aclFormat::ACL_FORMAT_ND,
shape.data(), shape.size(), *deviceAddr);
return 0;
}
constexpr uint64_t kBufAlignSize = 512;
inline uint64_t AlignUp(uint64_t num, uint64_t align) {
return ((num + align - 1) / align) * align;
}
#pragma pack(push, 1)
struct FfnDataDesc {
volatile int32_t flag;
volatile int32_t layer_id;
volatile int32_t expert_ids[0];
};
struct ScheduleContext {
struct CommonArea {
uint32_t session_num; // Number of attention nodes
uint32_t micro_batch_num;
uint32_t micro_batch_size;
uint32_t selected_expert_num; // topK + 1
uint32_t expert_num; // experts per layer
uint32_t attn_to_ffn_token_size;
uint32_t ffn_to_attn_token_size;
int32_t schedule_mode; // 0: Ffn only 1: Attention only
int8_t reserve0[96];
};
struct ControlArea {
int32_t run_flag; // 0 : exited 1 : running
int8_t reserve2[124];
};
struct FfnArea {
uint64_t token_info_buf; // Points to device memory.
uint64_t token_info_buf_size;
uint64_t token_data_buf; // Points to device memory.
uint64_t token_data_buf_size;
uint64_t polling_index;
int8_t reserve3[88];
uint64_t layer_ids_buf;
uint64_t layer_ids_buf_size;
uint64_t session_ids_buf;
uint64_t session_ids_buf_size;
uint64_t micro_batch_ids_buf;
uint64_t micro_batch_ids_buf_size;
uint64_t expert_ids_buf;
uint64_t expert_ids_buf_size;
uint32_t out_num;
int8_t reserve4[60];
};
struct AttentionArea {
uint64_t token_info_buf;
uint64_t token_info_buf_size;
uint64_t token_data_buf;
uint64_t token_data_buf_size;
uint32_t micro_batch_id;
int8_t reserve5[92];
};
CommonArea common;
ControlArea control;
AttentionArea attention;
FfnArea ffn;
int8_t reserve6[384]; // Padding to 1024 bytes.
};
static_assert(sizeof(ScheduleContext) == 1024, "ScheduleContext size must be 1024 bytes");
#pragma pack(pop)
int main() {
int32_t deviceId = 0;
aclrtStream stream;
auto ret = Init(deviceId, &stream);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("Init acl failed. ERROR: %d\n", ret); return ret);
// 2. 构造输入与输出,需要根据API的接口自定义构造
ScheduleContext scheduleContext = {};
scheduleContext.common.session_num = 1;
scheduleContext.common.micro_batch_num = 1;
scheduleContext.common.micro_batch_size = 8;
scheduleContext.common.selected_expert_num = 9; // topK + 1
scheduleContext.common.expert_num = 8;
scheduleContext.common.attn_to_ffn_token_size = 512;
scheduleContext.common.ffn_to_attn_token_size = 512;
scheduleContext.common.schedule_mode = 0; // Ffn only
scheduleContext.control.run_flag = 1; // running
scheduleContext.ffn.polling_index = 0;
// 标量参数
int64_t expertNum = 8; // 每层专家数 × layer_num
int64_t A = scheduleContext.common.session_num; // Attention worker数量
int64_t BS = scheduleContext.common.micro_batch_size; // micro batch size
int64_t K = scheduleContext.common.selected_expert_num;// topK + 1
int64_t H = 4096; // hidden size
int64_t Y = A * BS * K;
std::vector<int64_t> maxOutShapeValue = {A, BS, K, H};
int64_t tokenDtype = 0; // FP16
int64_t needSchedule = 0; // 仅做batching
int64_t layerNum = 1;
// 初始化Ffn token_info_buf
uint64_t flagAndLayerIdSize = sizeof(FfnDataDesc);
uint64_t tokenInfoSize = (sizeof(int32_t) * static_cast<uint64_t>(K) * BS + flagAndLayerIdSize) *
static_cast<uint64_t>(scheduleContext.common.micro_batch_num) *
static_cast<uint64_t>(scheduleContext.common.session_num);
scheduleContext.ffn.token_info_buf_size = tokenInfoSize;
void* tokenInfoBuf = nullptr;
ret = aclrtMalloc(&tokenInfoBuf, tokenInfoSize, ACL_MEM_MALLOC_HUGE_FIRST);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("malloc token info buf failed. ERROR: %d\n", ret); return ret);
scheduleContext.ffn.token_info_buf = reinterpret_cast<uint64_t>(tokenInfoBuf);
// 填充token_info_buf:flag=1,layer_id,expert_ids
std::vector<int32_t> hostTokenInfo(tokenInfoSize / sizeof(int32_t), 0);
int32_t* pInt = hostTokenInfo.data();
for (uint32_t s = 0; s < scheduleContext.common.session_num; ++s) {
for (uint32_t m = 0; m < scheduleContext.common.micro_batch_num; ++m) {
*pInt++ = 1; // flag
*pInt++ = 0; // layer_id
for (uint32_t i = 0; i < BS * K; ++i) {
*pInt++ = static_cast<int32_t>(i % scheduleContext.common.expert_num); // expert_id
}
}
}
ret = aclrtMemcpy(tokenInfoBuf, tokenInfoSize, hostTokenInfo.data(), tokenInfoSize, ACL_MEMCPY_HOST_TO_DEVICE);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("cpy token info buf failed. ERROR: %d\n", ret); return ret);
// 初始化Ffn token_data_buf
uint64_t tokenDataSize = static_cast<uint64_t>(Y) * H * sizeof(int16_t);
scheduleContext.ffn.token_data_buf_size = tokenDataSize;
void* tokenDataBuf = nullptr;
ret = aclrtMalloc(&tokenDataBuf, tokenDataSize, ACL_MEM_MALLOC_HUGE_FIRST);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("malloc token data buf failed. ERROR: %d\n", ret); return ret);
scheduleContext.ffn.token_data_buf = reinterpret_cast<uint64_t>(tokenDataBuf);
std::vector<int16_t> hostTokenData(Y * H, 1);
ret = aclrtMemcpy(tokenDataBuf, tokenDataSize, hostTokenData.data(), tokenDataSize, ACL_MEMCPY_HOST_TO_DEVICE);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("cpy token data buf failed. ERROR: %d\n", ret); return ret);
// 创建scheduleContext aclTensor
std::vector<int64_t> scheduleContextShape = {1024};
void* scheduleContextDeviceAddr = nullptr;
aclTensor* scheduleContextRef = nullptr;
std::vector<int8_t> hostCtxData(1024, 0);
std::memcpy(hostCtxData.data(), &scheduleContext, sizeof(ScheduleContext));
ret = CreateAclTensor(hostCtxData, scheduleContextShape, &scheduleContextDeviceAddr, aclDataType::ACL_INT8,
&scheduleContextRef);
CHECK_RET(ret == ACL_SUCCESS, return ret);
// 创建输出 aclTensor
std::vector<int64_t> yShape = {Y, H};
void* yDeviceAddr = nullptr;
aclTensor* yRef = nullptr;
ret = CreateAclTensorNoData(yShape, &yDeviceAddr, aclDataType::ACL_FLOAT16, &yRef);
CHECK_RET(ret == ACL_SUCCESS, return ret);
std::vector<int64_t> groupListShape = {expertNum, 2};
void* groupListDeviceAddr = nullptr;
aclTensor* groupListRef = nullptr;
ret = CreateAclTensorNoData(groupListShape, &groupListDeviceAddr, aclDataType::ACL_INT64, &groupListRef);
CHECK_RET(ret == ACL_SUCCESS, return ret);
std::vector<int64_t> oneDimShape = {Y};
void* sessionIdsAddr = nullptr;
aclTensor* sessionIdsRef = nullptr;
ret = CreateAclTensorNoData(oneDimShape, &sessionIdsAddr, aclDataType::ACL_INT32, &sessionIdsRef);
CHECK_RET(ret == ACL_SUCCESS, return ret);
void* microBatchIdsAddr = nullptr;
aclTensor* microBatchIdsRef = nullptr;
ret = CreateAclTensorNoData(oneDimShape, µBatchIdsAddr, aclDataType::ACL_INT32, µBatchIdsRef);
CHECK_RET(ret == ACL_SUCCESS, return ret);
void* tokenIdsAddr = nullptr;
aclTensor* tokenIdsRef = nullptr;
ret = CreateAclTensorNoData(oneDimShape, &tokenIdsAddr, aclDataType::ACL_INT32, &tokenIdsRef);
CHECK_RET(ret == ACL_SUCCESS, return ret);
void* expertOffsetsAddr = nullptr;
aclTensor* expertOffsetsRef = nullptr;
ret = CreateAclTensorNoData(oneDimShape, &expertOffsetsAddr, aclDataType::ACL_INT32, &expertOffsetsRef);
CHECK_RET(ret == ACL_SUCCESS, return ret);
void* dynamicScaleAddr = nullptr;
aclTensor* dynamicScaleRef = nullptr;
ret = CreateAclTensorNoData(oneDimShape, &dynamicScaleAddr, aclDataType::ACL_FLOAT, &dynamicScaleRef);
CHECK_RET(ret == ACL_SUCCESS, return ret);
std::vector<int64_t> actualTokenNumShape = {1};
void* actualTokenNumAddr = nullptr;
aclTensor* actualTokenNumRef = nullptr;
ret = CreateAclTensorNoData(actualTokenNumShape, &actualTokenNumAddr, aclDataType::ACL_INT64, &actualTokenNumRef);
CHECK_RET(ret == ACL_SUCCESS, return ret);
// 创建maxOutShape aclIntArray
aclIntArray* maxOutShapeArray = aclCreateIntArray(maxOutShapeValue.data(), maxOutShapeValue.size());
// 3. 调用CANN算子库API
uint64_t workspaceSize = 0;
aclOpExecutor* executor = nullptr;
ret = aclnnFfnWorkerBatchingGetWorkspaceSize(scheduleContextRef, expertNum, maxOutShapeArray, tokenDtype,
needSchedule, layerNum, yRef, groupListRef, sessionIdsRef,
microBatchIdsRef, tokenIdsRef, expertOffsetsRef, dynamicScaleRef,
actualTokenNumRef, &workspaceSize, &executor);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclnnFfnWorkerBatchingGetWorkspaceSize failed. ERROR: %d\n", ret); return ret);
void* workspaceAddr = nullptr;
if (workspaceSize > 0) {
ret = aclrtMalloc(&workspaceAddr, workspaceSize, ACL_MEM_MALLOC_HUGE_FIRST);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("allocate workspace failed. ERROR: %d\n", ret); return ret);
}
ret = aclnnFfnWorkerBatching(workspaceAddr, workspaceSize, executor, stream);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclnnFfnWorkerBatching failed. ERROR: %d\n", ret); return ret);
// 4.(固定写法)同步等待任务执行结束
ret = aclrtSynchronizeStream(stream);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtSynchronizeStream failed. ERROR: %d\n", ret); return ret);
// 5. 获取输出结果,将device侧内存上的结果拷贝至host侧
int64_t actualTokenNum = 0;
ret = aclrtMemcpy(&actualTokenNum, sizeof(int64_t), actualTokenNumAddr, sizeof(int64_t), ACL_MEMCPY_DEVICE_TO_HOST);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("copy actual_token_num failed. ERROR: %d\n", ret); return ret);
LOG_PRINT("actual_token_num = %ld.\n", actualTokenNum);
std::vector<int64_t> groupListHost(expertNum * 2, 0);
ret = aclrtMemcpy(groupListHost.data(), expertNum * 2 * sizeof(int64_t), groupListDeviceAddr,
expertNum * 2 * sizeof(int64_t), ACL_MEMCPY_DEVICE_TO_HOST);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("copy group_list failed. ERROR: %d\n", ret); return ret);
for (int64_t i = 0; i < expertNum; i++) {
LOG_PRINT("group_list[%ld] = [expert_id=%ld, token_num=%ld]\n", i, groupListHost[i * 2], groupListHost[i * 2 + 1]);
}
// 6. 释放aclTensor与aclIntArray
aclDestroyTensor(scheduleContextRef);
aclDestroyTensor(yRef);
aclDestroyTensor(groupListRef);
aclDestroyTensor(sessionIdsRef);
aclDestroyTensor(microBatchIdsRef);
aclDestroyTensor(tokenIdsRef);
aclDestroyTensor(expertOffsetsRef);
aclDestroyTensor(dynamicScaleRef);
aclDestroyTensor(actualTokenNumRef);
aclDestroyIntArray(maxOutShapeArray);
// 7. 释放device资源
aclrtFree(scheduleContextDeviceAddr);
aclrtFree(yDeviceAddr);
aclrtFree(groupListDeviceAddr);
aclrtFree(sessionIdsAddr);
aclrtFree(microBatchIdsAddr);
aclrtFree(tokenIdsAddr);
aclrtFree(expertOffsetsAddr);
aclrtFree(dynamicScaleAddr);
aclrtFree(actualTokenNumAddr);
aclrtFree(tokenInfoBuf);
aclrtFree(tokenDataBuf);
if (workspaceSize > 0) {
aclrtFree(workspaceAddr);
}
Finalize(deviceId, stream);
return 0;
}