GLM 5.3 Flash Max 搭配 OpenCode 用时最短(20 分钟)且得分最高(96.89%),Astra 6.0 Max 搭配 Codex 耗时最长(37 分钟)。 测试记录了各组合的耗时、token 用量与工具调用成功率。
我一直在使用不同模型和框架组合测试一个简单的提示,以找出哪一个能产生最佳效果。我在 /goal 模式下执行此操作。
提示:构建一个单页的 Three.js 科幻机库,包括悬浮无人机、动画警告灯、发光跑道条带和微妙的体积风格雾平面。包括无人机编队切换和电影感摄像机路径。输出一个自包含的 HTML 文件,并内联 JavaScript。
所有 GLM 运行都标记为 GLM 5.3 Flash Max;输入 token 包括缓存的输入。输出 token 是生成的总数,包括推理;当框架单独报告推理时,推理列显示该子集。DSH 适配器不报告单独的推理计数。DSH 持续时间是两次轮次的活动时间之和,不包括轮次之间的暂停。工具错误记录失败的工具事件。破折号表示不可用或未报告。
I've been testing a simple prompt with different model and harness combinations to work out which one produces best results. I do this in /goal mode.
Prompt: Build a single-page Three.js sci-fi hangar with hovering drones, animated warning lights, emissive runway strips, and subtle volumetric-style fog planes. Include drone formation toggle and cinematic camera path. Output one self-contained HTML file with inline JavaScript.
All GLM runs are labelled GLM 5.3 Flash Max ; input tokens include cached input. Output tokens are the generated total, including reasoning; when a harness reports reasoning separately, the reasoning column shows that subset. The DSH adapter does not report a separate reasoning count. DSH durations sum active turn time across both turns, excluding the pause between turns. Tool errors are recorded failed tool events. A dash means unavailable or not reported.