Tencent paper reveals non-thinking mode increases response failures by up to 48% in multimodal AI models

51 minutes ago 1



Tencent researchers have quantified something that heavy users of multimodal AI models have likely noticed anecdotally: when these systems skip the “thinking” step and jump straight to answers, things break a lot more often. The gap between thinking and non-thinking inference failure rates reaches as high as 48.64% in flagship models, according to the team’s new paper published on arXiv. The study, titled “Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs,” introduces a diagnostic benchmark called PatternEval. It’s designed to catch what traditional accuracy metrics miss entirely. A model can score well on correctness while simultaneously producing responses riddled with contradictions, repetitions, and reasoning that looks sophisticated but goes nowhere. What PatternEval actually measures PatternEval consists of 2,415 multimodal prompts spread across multiple task categories. Rather than simply asking “did the model get the right answer,” the benchmark evaluates the quality and coherence of the response itself. The researchers identified four dominant failure patterns that plague non-thinking outputs. Chain-of-thought leakage occurs when fra...

Read Entire Article