I agree it’s a great step. But the deepseek models also don’t perform to the same level of fable/sol. If we optimize/finetune to deepseek traces, wouldn’t it be suboptimal? What would the benefit be?
You let the smarter model explore the traces and figure out where the current harness' bottlenecks are for the current LLM. Then you can adjust prompts or tools to fix those.