``` models: llm: model: ai/gemma4:e2b runtime_flags: - "--reasoning-budget" - "0" ``` When using this configuration on MacOS, the model outputs its reasoning inside the answers's content ie: "The user wants ...."
When using this configuration on MacOS, the model outputs its reasoning inside the answers's content ie: "The user wants ...."