Meta's highest-paid employee Alexandr Wang can't stop making fun of Google

Meta's highest-paid employee Alexandr Wang can't stop making fun of Google

Alexandr Wang, Meta’s chief AI officer and highest-paid employee recently took a jab at Google’s Gemini after a benchmark data shoed Meta’s latest AU model climbing near the very top of the industry rankings. In a post shared on social media platform X (formerly known as Twitter) Wang wrote, “I really hate to say it, but… gemini who? 🚀💨” alongside a performance chart showing Meta’s Muse Spark 1.3 surpassing several Gemini variants on the Artificial Analysis Intelligence Index. According to Artificial Analysis, Muse Spark 1.3 scored 62 points, placing it behind only Anthropic’s Claude Fable 5.1 (66) and Claude Opus 5 (63). The model outperformed Google’s Gemini 1.5 Pro and Gemini 1.5 Flash, which scored 61 and 53 respectively.The benchmark behind Alexandr Wang’s dig at GoogleWang's comment was posted in response to a thread from Artificial Analysis announcing that Meta had released Muse Spark 1.3, the company's fourth Muse Spark model release in just five months. According to the benchmark firm, Muse Spark 1.3 (max) — currently in limited preview for Meta's partners — scored 62 on the Artificial Analysis Intelligence Index, trailing only Claude Fable 5.1 and Claude Opus 5 in their top configurations.The publicly available variant, Muse Spark 1.3 (xhigh), scored 61 on the same index, putting it in a tie with GPT-5.6 Sol (max), Grok 4.6 (high), and Claude Opus 5 (high). That score represents a 4-point jump from Muse Spark 1.2's score of 57 in August, and an 8-point gain over Muse Spark 1.1's score of 53 in July. It landed just behind Claude Fable 5.1 (max, 66), Claude Opus 5 (max, 63), and Claude Fable 5 (max, 62).Where the gains came fromAccording to Artificial Analysis, both new Muse Spark variants improved primarily in two areas: agentic task performance and scientific reasoning.On agentic work, Muse Spark 1.3 (xhigh) posted a 12-point gain over Muse Spark 1.2 on the Tau3-Bench Banking evaluation, rising from 35% to 47%, along with a 5-point gain on Terminal-Bench 2.1 (80% to 85%) and a jump in GDPval-AA v2 Elo rating from 1,615 to 1,709. The limited-preview Muse Spark 1.3 (max) variant pushed even further, reaching 52% on Tau3-Bench Banking — the highest score of any model on that specific benchmark — and a GDPval-AA v2 Elo of 1,754. Artificial Analysis noted that Muse Spark 1.3 (max) achieves these higher scores in part by using significantly more reasoning tokens than the xhigh variant, reasoning 62% more on GDPval-AA v2 and 28% more on Tau3-Bench Banking.On scientific reasoning, the standout gain came on CritPt, where the xhigh variant improved 8 points over Muse Spark 1.2, from 18% to 26%. GPQA Diamond rose 4 points, from 90% to 94%, while Humanity's Last Exam and SciCode each posted more modest 2-to-3-point gains.The cost advantageBeyond performance, Muse Spark 1.3 also demonstrated efficiency. The xhigh variant costs about $0.55 per task, significantly lower than GPT‑5.6 Sol and Grok 4.6, which cost nearly double. Analysts noted this places Muse Spark on the “Pareto frontier” for intelligence versus cost, making it one of the most competitive models in terms of value.For those unaware, Meta has released four Muse Spark models in just five months, underscoring its aggressive push in frontier AI.

Original Source

Read the full article at Timesofindia →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.