分类筛选
找到 1605 个 Demo
Mistral didn't stop at Le Chaton Fat. They droppe...
Mistral didn't stop at Le Chaton Fat. They dropped Le Gros Chaton. 30T Sparse MoE. 256 Experts. 1M context window. Beats Claude 3.7 Sonnet, Gemini 2.5 Pro, and GPT-4.1 on SWE-bench and long context. And it's open weights. Under the Mistral Research License. The French are abso https://t.co/IqiOxMuyWm
AI Models and Their Release Year GPT-1 - 2018 BER...
AI Models and Their Release Year GPT-1 - 2018 BERT - 2018 GPT-2 - 2019 GPT-3 - 2020 DALL-E - 2021 GitHub Copilot - 2021 Midjourney - 2022 Stable Diffusion - 2022 Whisper - 2022 ChatGPT (GPT-3.5) - 2022 LLaMA - 2023 GPT-4 - 2023 Claude - 2023 Llama 2 - 2023 Gemini 1.0 - 2023 Sora - 2024 Claude 3 - 2024 Gemini 1.5 - 2024 Llama 3 - 2024 GPT-4o - 2024 Claude 3.5 - 2024 o1 - 2024 DeepSeek V3 - 2024 Gemini 2.0 - 2024 DeepSeek R1 - 2025 o3 - 2025 GPT-4.5 - 2025 Gemini 2.5 - 2025 Llama 4 - 2025 Claude 4 - 2025 GPT-5 - 2025 Gemini 3 - 2026 Claude 4.8 - 2026 o4 - 2026 DeepSeek R2 - 2026 Which AI model changed everything for you?
6/16(火)90分ジョグ 🗒️20.01km 🔧Free 👟ディヴィエイトピュアニトロ 明...
6/16(火)90分ジョグ 🗒️20.01km 🔧Free 👟ディヴィエイトピュアニトロ 明日のポイント練に向け淡々といい感じに走れた 明日は練習会前に個人練行う予定 週末のレースに向けいい形で走りきりたい #GeminiRunners #Gemini3 #まるお製作所RC https://t.co/4eG1tOZufG
Great to see the breadth of things that can be bui...
Great to see the breadth of things that can be built with Gemini and Veo. Come try yourself @ https://t.co/vEPXhE2Hkk
Today, we’re continuing to push the boundaries of...
Today, we’re continuing to push the boundaries of AI with our release of Gemini 3.1 Pro. This updated model scores 77.1% on ARC-AGI-2, more than double the reasoning performance of its predecessor, Gemini 3 Pro. Check out the visible improvement in this side-by-side comparison, showing Gemini 3.1 Pro’s crisp animation built with pure code. Read more about today’s 3.1 Pro update: https://t.co/vABdcMSE3f
StepFun AI Releases Step-Audio-R1: A New Audio LLM...
StepFun AI Releases Step-Audio-R1: A New Audio LLM that Finally Benefits from Test Time Compute Scaling StepFun’s Step-Audio-R1 is an open audio reasoning LLM built on Qwen2 audio and Qwen2.5 32B that uses Modality Grounded Reasoning Distillation and Reinforcement Learning with Verified Rewards to turn long chain of thought from a liability into an accuracy gain, surpassing Gemini 2.5 Pro and approaching Gemini 3 Pro on comprehensive audio benchmarks across speech, environmental sound and music while providing a reproducible training recipe and vLLM based deployment for real world audio applications..... Full analysis: https://t.co/teIjWvM4xV Paper: https://t.co/6Mihzrkqvq Project: https://t.co/ItRvvAWdZr Repo: https://t.co/9hnrGeP0lG Model weights: https://t.co/T8IEBpNC2b @StepFun_ai
Everyone is talking about 's 3.0, but what is so...
Everyone is talking about @GoogleAI's @GeminiApp 3.0, but what is so special about it? Here is a quick breakdown of the key differences between Gemini 2.5 and the new Gemini 3: Reasoning Capabilities (The "Deep Think" Leap) Gemini 2.5 introduced "thinking models" (like Gemini 2.5 Flash) that could pause to reason before responding. It was a major step forward in accuracy. Gemini 3 introduces "Deep Think", a far more advanced reasoning mode. It allows the model to handle PhD-level science, complex math, and nuanced strategic planning much better than 2.5. (Early benchmarks show Gemini 3 acts more like a thoughtful partner than just a rapid-response engine.) "Vibe Coding" and Agentic Workflows Gemini 2.5 was a strong coding assistant, decent for generating snippets and debugging. Gemini 3 is designed for "Vibe Coding"... meaning you can use natural language ("just make it look cool and floaty") and the model understands the intent behind the code. It is also built for Agentic Workflows (via the new Google Antigravity platform), allowing it to act autonomously to plan, execute, and complete multi-step software engineering tasks rather than just writing code blocks. Early benchmarks show Gemini 3 Pro has shown a 50% improvement over Gemini 2.5 Pro in resolving complex software engineering challenges. Multimodal Understanding Gemini 2.5 had excellent video and image processing capabilities (1M+ context window). Gemini 3 on the other hand is touted as the "best model in the world for multimodal understanding." It has significantly improved spatial reasoning and visual comprehension. For example, it can better understand the physical layout of a room from a video or untangle complex handwritten notes in multiple languages. Performance and Benchmarks Gemini 2.5 was a leaderboard topper upon release. Gemini 3 along similar lines, has immediately taken the #1 spot on the LMArena Leaderboard (with a score of 1501 Elo) and holds the record on "Humanity's Last Exam," a benchmark for general expertise. Anything else that I missed?
Built this Anime game site with google AI studio,...
Built this Anime game site with google AI studio, what do you think? sorry about the background sound, i was watching anime lol https://t.co/vrAWef9gBb
大人向けに学習した音声AIは、子どもの前で失速する。それを体系的に示したベンチマーク「ChildVo...
大人向けに学習した音声AIは、子どもの前で失速する。それを体系的に示したベンチマーク「ChildVox」が公開された(https://arxiv[.]org/abs/2605.29257)。 USC・OSU・UCLAなど複数大学の共同研究。生後すぐの心音・呼吸音から、乳児の泣き声・笑い声、幼児の「だだ」「ばば」のような正準音節(子音と母音を組み合わせた最初の音節)、就学後の会話まで。子どもの発達軌跡を丸ごとカバーした統合ベンチマークで、17データセット・20以上のタスクを一つの評価基盤に統合している。 面白い結果がいくつかある。まず「1つのモデルが全タスクを制する」ことはなかった。心音・呼吸音ではSSAST・WavLM-Largeが優位で、音声認識ではWhisper-Largeが最も精度が高く(子ども音声MyST WER 14.80%)、発声の感情分類ではQwen2-Audioが健闘する。モデルが捉える音の「種類」が根本的に違うため、タスクで最適なモデルが変わる。 特に目を引くのがASD(自閉症スペクトラム)の子どもへの認識精度。Whisper-Largeは一般的な子ども音声(MyST)でWER 14.80%だが、ASD評価データ(ADOS)では同じモデルが40.20%まで悪化する。約2.7倍の劣化で、発話特性の多様さが現状モデルにとっていかに難しいかを示している。 ゼロショット(タスク固有の学習なし)のGemini 2.5/3.5 Flashは、ChildVoxで専門ファインチューニングされたモデルに全タスクで負けた。汎用大型モデルでも勝てない領域がまだある。さらにAudioFlamingo3という大規模音声言語モデルは、ラベルを答えるよう指示しているのに「赤ちゃんが笑っているように聞こえます」と自由記述を返すなど指示追従の失敗が目立った。数字よりこの失敗事例の方がモデルの限界をリアルに示している。
LLMエージェントの「記憶」を、使うたびにグラフ(ノードとエッジの繋がり)の構造自体が変わる仕組みと...
LLMエージェントの「記憶」を、使うたびにグラフ(ノードとエッジの繋がり)の構造自体が変わる仕組みとして設計したFluxMemが発表された(https://arxiv[.]org/html/2605.28773v1)。 従来のメモリ拡張エージェントは保存形式と検索方法が決め打ちで、実行中のフィードバックに応じて「繋ぎ方」を変えられない。これは2種類の失敗を生む。必要な文脈を取り逃す「Under-connection(接続不足)」と、無関係な情報が混入して幻覚を起こす「Over-connection(接続過剰)」だ。FluxMemはこれを認知科学から着想を得て解決した。脳が経験に応じてシナプス結合を強化・剪定するように、記憶グラフも実行しながら進化させる。 グラフは3層で構成される。事実・ドキュメントを蓄える意味知識層、具体的な行動履歴を記録するエピソード層、過去の成功パターンを圧縮した手順層(再利用可能なスキル)だ。 進化は3ステージで進む。タスク開始時に3層から関連情報を検索して初期グラフを組む(Stage I)。実行フィードバックを見てリアルタイムでグラフを書き換え、リンクの追加・削除から情報の粒度の調整まで行う(Stage II)。タスク完了後は成功軌跡を類似ケースでグループ化して手順スキルに蒸留し、PEMS(成熟度を評価する独自スコア)で収束を判定する。成熟したスキルは次のタスクで検索をスキップして直接発動できる(Stage III)。 3ベンチマークで全SOTA。LoCoMo(長文推論)はGPT-4.1-miniで95.06(LLM採点評価。Full Context 81.23、次点のEverMemOS 93.05を超えた)。Mind2Web(ウェブ操作、手動フィルタなしの現実設定)のCross-Task成功率はGPT-4.1-miniで8.1%(AWM 3.6%の2倍超)、Gemini-2.5-flashで9.6%(AWM 5.6%)。GAIA(汎用タスク)はGPT-5-miniで平均76.36%(Langfun 71.52%、Alita 72.73%を上回る)、Kimi K2では64.85%(Flash-Searcher 52.12%から+12.73pt)。 アブレーション(各ステージを外した検証)の結果が興味深い。事実検索型のLoCoMoではStage IIが最重要で、除去するとGPT-4.1-miniで95.06→85.32に低下する。複雑推論が必要なMind2WebではStage IIIが主効果で、除去するとCross-Task成功率が8.1→3.2に急落する。「どの記憶操作が効くか」はタスクの性質次第という設計論的な知見として面白い。 コード: https://github[.]com/zjunlp/LightMem
o3 and Gemini 2.5 Pro both failed. This is next A...
o3 and Gemini 2.5 Pro both failed. This is next AGI test. https://t.co/2WyIkYV6lZ
Creati just went WILD with Nano Banana Pro 🍌🔥 O...
Creati just went WILD with Nano Banana Pro 🍌🔥 On Nov 21, Higgsfield opened FREE access to Google’s Nano Banana Pro Image model — zero cost, full power. Black Friday officially UPGRADDED ⚡ #NanoBanana #Higgsfield #Gemini3 #BlackFriday https://t.co/ptjZMX7vTO