I’m not sure why there’s so much focus on comparing Qwen’s 27B models with older Opus versions. 3.6 was pretty comparable to Sonnet 4.6 for coding, and now 3.8 gets it to around Sonnet 5 level (SWE Bench Pro 61.7 vs. S5’s 63.2). That’s pretty a awesome achievement and I think comparing it to Sonnet sets more realistic expectations.
I’m not sure why there’s so much focus on comparing Qwen’s 27B models with older Opus versions. 3.6 was pretty comparable to Sonnet 4.6 for coding, and now 3.8 gets it to around Sonnet 5 level (SWE Bench Pro 61.7 vs. S5’s 63.2). That’s pretty a awesome achievement and I think comparing it to Sonnet sets more realistic expectations.