BETA nonprofit public democratic european moderated

Search

#benchmarks
marcus 4d

MOONSHOT ACABA DE LANZAR EL MODELO OPEN SOURCE MÁS GRANDE DE LA HISTORIA se llama Kimi K3. 2.8 billones de parámetros. ventana de contexto de 1 millón de tokens. multimodal nativo. lo interesante no es solo el tamaño. usa dos arquitecturas nuevas: Kimi Delta Attention y Attention Residuals. → decodificación hasta 6.3x más rápida en contextos de 1M tokens → ~25% más eficiencia de entrenamiento con menos de un 2% de coste extra → pensado para coding de largo horizonte y trabajo agéntico según los propios benchmarks de Moonshot, K3 solo queda por detrás de Claude Fable 5 y GPT-5.6 Sol. y por delante de Claude Opus 4.8. todo open source. #OpenSource #AI #MachineLearning

The reform of the ETS system is absolutely necessary for the reconstruction of the competitiveness of the Polish and European industry. I will be dealing with one of the key regulations in this area. A proposal from @EU_Commission will be published on Friday, and we are getting started. #industry #emissions #ETS #benchmarks (translated)

Ich finde, es sollte in Deutschland eine Möglichkeit geben, die Kinder neben den klassischen Schulen zu Hause zu unterrichten. Wenn sie dann staatliche Tests absolvieren und die geforderten Leistungen regelmäßig erreichen, sollten die Eltern pro Kind 12.000 € im Jahr steuerfrei erhalten. Der Staat gibt laut Statistischem Bundesamt sowieso diese Summe pro Schüler aus. Warum also nicht den Eltern überlassen, wie sie das Geld einsetzen, solange das Kind die Benchmarks erreicht? #Bildung #Homeschooling #Steuern

antirez Jun 17

OpenAI may delay GPT6 (or even 5.6) before making sure could not be blocked like Fable. Or they could play it smart, publishing only the benchmarks that show the improvements on certain area, providing a very censored model in the cyber-security side, and cross their fingers. #OpenAI #GPT6 #CyberSecurity #Nospecificcountrycodesarepresentinthetweet.

🇨🇳🇺🇸 The most important AI leaderboard today isn't benchmarks. It's adoption. #OpenRouter's rankings show a striking trend: Chinese models are rapidly overtaking US models in real-world developer usage ⚠ Models from DeepSeek, Kimi, MiniMax, GLM, and Qwen now account for a large https://t.co/H2dYoBiWQ4 #AI #China #TechTrends #CN #US

antirez May 28

Anthropic did a big strategic error. Normally they compare their models with their old models. Instead today, now that everybody knows how strong GPT 5.5 is at coding, they put it in the mix, basically showing all their customers that the benchmarks can't be trusted. https://t.co/up73bHAfen #AI #Anthropic #GPT5.5 #Therearenospecificcountrycodesrelatedtothecontentofthetweet.

Marta Kos May 26

A big step forward for Albania on its path to EU membership! With Albanian Prime Minister @ediramaal and Cypriot Deputy Minister for European Affairs @marilena_raouna in Brussels today, Albania reached the interim benchmarks, and we set the closing benchmarks for fundamentals https://t.co/oyf8mftyn7 #Albania #EUMembership #BrusselsSummit #AL #CY

Yann LeCun May 14

RT @logic_int: Aleph, our fully autonomous AI agent system for formal verification, aced all major theorem proving benchmarks including Put… #AI #AutonomousSystems #FormalVerification #Nospecificcountriesarementionedinthetweet.