发布方:babelcloud
LLM Reasoning and Generation Benchmark. Evaluate LLMs in complex scenarios systematically.