A benchmark to evaluate language models on questions I've previously asked them to solve.
一个评估语言模型在回答我之前提出的问题时的性能基准。【此简介由AI生成】