# 增强式学习# LLM# 大模型在slime上用Search-R1训练Qwen2.5-3B搜索智能体在 NQ 和 HotpotQA 上训练与评估 slime Search-R1 训练后的 Qwen2.5-3B 模型