分类筛选
找到 1605 个 Demo
I really like this usage dashboard in Google AI St...
I really like this usage dashboard in Google AI Studio. Had some free time today and was using Hermes. Agent with Gemini 3.5 Flash and Gemini 3.1 Flash Lite. Gemini 3.5 Flash is really good in Hermes. I am very new to Hermes and trying to figure out how I can use it to automate parts of my daily work. It has already helped a lot with my Notion and streamlining the mess I had built up over the years.
Google really went from Bard to Gemini 2.5 Pro.
Google really went from Bard to Gemini 2.5 Pro. https://t.co/fhVNoTMyt8
「ローカルLLM=安い」という前提を壊してくれる記事。どのようなケースでローカルLLMを使うべきかが...
「ローカルLLM=安い」という前提を壊してくれる記事。どのようなケースでローカルLLMを使うべきかが、わかりやすく整理されている。 ・損保ジャパンDXのエンジニアによる「ローカルLLMをいつ使うべきか」の実践レポート ・「ローカルLLM=コストが安い」という一般的な認識を否定する内容 ・コスト面での損益分岐点は月約110億トークン(1日5億トークン規模)と極めて高く、ほとんどの場合はAPI利用の方が有利と結論 ・実例として、社内ナレッジ検索RAGシステムで「会話のトピック逸脱検知」と「追加検索判定」という2つの分類タスクをローカルモデルに切り出した ・素のGemma 4 E2B(2Bモデル)でのドリフト検知率は61% ・LoRAでファインチューニング後は97%に向上。Gemini-2.5-flash(89%)や親モデルの31B版(88%)を上回る成績 ・特化のための手順は3つ: ①大規模モデル(Gemma 4 31B)で教師データを生成、②別モデル(Claude Sonnet 4.6)で評価データを作成し評価リークを防止、③品質ゲートで不合格サンプルを除外 ・ローカルLLMが優位な軸は「タスク特化での精度」「レイテンシの安定性」「ガバナンス」の3つ ・レイテンシ安定性: p95が2.15秒でエラー率0%という予測可能な性能を確保 ・ガバナンス面では医療・金融など規制産業でデータを外部に出せない場合の必須選択肢として位置づけ ・推奨ユースケースは「規制産業での定型的な判定・分類」「エージェントループ内の高頻度軽量ステップ」「低レイテンシ/オフライン対応が必須の現場」 ・最も現実的な構成はハイブリッド(7B〜13Bクラスとフロンティアモデルの組み合わせ)とまとめている https://t.co/umtdBzVtSG
🔥 VibeThinker-3B is a 3B open-source (MIT) reason...
🔥 VibeThinker-3B is a 3B open-source (MIT) reasoning model that reaches the band of systems hundreds of times larger on verifiable math and code. Math: 94.3 on AIME26, 89.3 on HMMT25, 93.8 on BruMO25, 76.4 on IMO-AnswerBench. With CLR test-time scaling those rise to 97.1 / 95.4 / 99.2 / 80.6. Code: 80.2 Pass@1 on LiveCodeBench v6 and 38.6 on OJBench. Instruction following holds at 93.4 IFEval after the reasoning RL. Built on Qwen2.5-Coder-3B via the Spectrum-to-Signal pipeline: curriculum two-stage SFT with Diversity-Exploring Distillation → MGPO RL across math/code/STEM at a single 64K context → Long2Short Math RL → Offline Self-Distillation → Instruct RL. CLR samples K=32 trajectories, extracts M=5 decision-relevant claims, then self-verifies them into a nonlinear reliability score — adding accuracy with zero extra parameters. On unseen LeetCode contests (Apr 25–May 31), it passed 123/128 first-attempt Python submissions — 96.1% acceptance, near GPT-5.2 and Gemini 3 Flash 👀 The catch: on knowledge-heavy GPQA-Diamond it sits at 70.2 (72.9 with CLR), still trailing large models. The research team frames this as the Parametric Compression-Coverage Hypothesis — reasoning compresses into a small core, broad knowledge still needs scale. Full analysis: https://t.co/EUfaw5IFzE Paper: https://t.co/AdJ7qwPgks Model weight: https://t.co/W7qUTfL7PF Repo: https://t.co/7ZGq9klKWT
While millions of people keep mindlessly burning t...
While millions of people keep mindlessly burning thousands of dollars every year on subscriptions and building everything from scratch, a 20-year-old guy downloaded free open-source repositories from GitHub and, within a week, built an autonomous AI empire generating insane profits! Most developers spend months searching for useful code only to find outdated junk, because the level of noise in today’s open-source ecosystem has become critically high. The breakthrough in income and automation begins when you stop reinventing the wheel and start leveraging production-tested, institutional-grade tools that completely eliminate technical routine. With specialized behavioral guidelines and skill files, such as tools created by Andrej Karpathy and Matt Pocock, you can strictly control your AI agents, preventing them from overwriting neighboring code and forcing them to make only precise, targeted edits. Moreover, utilities like free-claude-code allow you to redirect expensive API requests to free alternative backends such as Google AI Studio, while data compression servers can reduce token costs by up to 95%. When working on complex projects, quantitative analysts and developers use code intelligence systems that transform entire repositories into smart dependency graphs, helping AI instantly understand folder structures and maintain rock-solid memory across different sessions. If your goal is rapid content production or business automation, repositories like MoneyPrinterTurbo can generate viral short-form videos with voiceovers and music from a single text prompt, while powerful parsers such as crawl4ai and Microsoft’s markitdown convert websites, Word documents, and Excel spreadsheets into clean Markdown for AI context within seconds. Instead of exhausting manual programming, you simply deploy this ready-made infrastructure on top of the free local interface open-webui in a single evening and launch an unlimited production pipeline that handles client tasks on autopilot.
Review of my AI predictions for 2025: Lab will de...
Review of my AI predictions for 2025: Lab will declare AGI -> ❓❌ - not a lab, but Sam Altman CEO of OpenAI said this: "my proposal is that we agree that you know AGI kinda went whooshing by. It didn't change the world that much, or it will in the long term, but okay, fine, we built AGIs." (https://t.co/OIk3jVXJYn) - overall i get the feeling that everyone has exactly done that, we have all quietly acknowledged that 5 years ago we would call current models AGI, but the goalposts shifted - it's debateable, but i don't count this as a win Lab mentions ASI -> ❓✅ - OpenAI: "In ten more years, I believe we are almost certain to build superintelligence" (https://t.co/RR6wJAjhRe) - Sam Altman says it's a "superintelligence research company" (https://t.co/f7OnzpttfK) - Zuck and Elon also made remarks on ASI, but overall this is just a bad prediction because it is too vague Model Fiesta in Q1 -> ✅ - mostly yes, but some expected models came only later in April - Google: Gemini 2.5 Pro Preview and Gemma 3 + bunch of Gemini checkpoints (Gemini 2.5 Flash Preview in April) - Anthropic: Sonnet 3.7 - OpenAI: o3-mini and GPT-4.5 (GPT-4.1 o3 and o4-mini in april) - Meta: (Llama-4 in April) - Qwen: qwq, qwq-max, qwen-2.5-max (qwen3 in april - Mistral: Mistral Small 3.1, Mistral Saba others: DeepSeek-R1, Kimi-K-1.5 - Overall pace of frontier model releases was definitely accelerated in 2025. China bros really carried open-source! Agents take off -> ✅ - probably the clearest win in my predictions with Claude 4.5 Opus, Claude Code, Codex, Gemini CLI and all the other agentic harnesses Computer use takes off -> ✅ - we saw custom computer use models from OpenAI, Google, Anthropic - OS-World went from 27.1% (uitars-72b-dpo) to 72.6% (Opus 4.5) with harness (https://t.co/slCdPCokZ1) - basically every lab has computer-use integrated into their chat experience Release of massive models like Claude 4, Gemini 3, GPT-5, Grok 4 -> ✅ - we finally got another Opus, GPT-4.5, Grok-4, Gemini 3 Pro and GPT-5 - notable mention: Kimi-K2 and rise of massive ultra-sparse MoE's Release of o3, o4, and o5 -> ✅ - o3 and o4 (https://t.co/Ab32VbRrN7) - o5 is tricky, but I count it, as this was more a prediction of model release pace - after o4-mini we saw at least 3 more iterations of reasoning models from OpenAI with GPT-5, GPT-5.1 and GPT-5.2 that introduced new techniques like context-compaction, higher reasoning efficiency, better reasoning allocation, xhigh compute settings and more. surely one of them would have been o5 Open-Source Replication of o3 -> ✅ - Kimi-K2 Thinking, DeepSeek-V3.2, GLM-4.7, MiMo-V2 are all better on Artificial Intelligence Index - on ECI o3 is rated 144-149, Kimi-K2 Thinking: 143-147, DeepSeek V3.2: 141-147 - GLM-4.7 would absolutely smash o3 in coding and agentic tasks Frontier Math > 80% -> ❌ - GPT-5.2 xhigh achieves 41% (https://t.co/7WcmUfKpGT) - simply overestimated progress on mathematics - i thought it's the easiest verifiable domain and progress would be much faster than anyone anticipates SWE-bench > 90% ->❌ - Claude 4.5 Opus 80.9%(https://t.co/HEpHBEw9ZX) - again overestimated progress in this verifiable domain ARC-AGI 2 >80% within 9 months -> ❓✅ - GPT-5.2 Pro achieves only 54.2%(https://t.co/bKh8xsHC7p) - BUT, they didn't go for the 1k-10k$/task compute budget like o3-preview, which I expected when I made this prediction - I think >80% is currently possible, as poetiq reached 75% withbasic scaffolding and a <10$/task budget (https://t.co/kVvbrsIwc2) - i count this as a win in my book 10+ million context models -> ❓❌ - Llama-4 was released with an advertised context window of 10 million tokens, but that is obviously not real, so I don't count this - but nontheless, we saw massive long context improvements on MRCR and went from around 30% to 93.8% @ 1 million context (https://t.co/1O656klUwa)
100+ AI Tools to replace your tedious work: 1. Re...
100+ AI Tools to replace your tedious work: 1. Research * ChatGPT * YouChat * Abacus * Perplexity * Copilot * Gemini 2. Image * Higgsfield AI Soul * GPT-4o * Midjourney * Grok 3. Productivity * Gamma * Grok 3 * Perplexity AI * Gemini 2.5 Flash 4. Writing * Jasper * Jenny AI * Textblaze * Quillbot 5. Video * Klap * Kling * InVideo * HeyGen * Runway 6. Meeting * Tldv * Otter * Noty AI * Fireflies 7. SEO * VidIQ * Seona AI * BlogSEO * Keywrds ai * Outrank AI 8. Presentation * Decktopus * Slides AI * Gamma AI * Designs AI * Beautiful AI 9. Design * Canva * Flair AI * Designify * Clipdrop * Autodraw * Magician design 10. Audio * Lovo ai * Eleven labs * Songburst AI * Adobe Podcast 11. Marketing * Pencil * Ai-Ads * AdCopy * Simplified * AdCreative 12. Startup * Tome * Ideas AI * Namelix * Pitchgrade * Validator AI 13. Social media management * Tapilo * Typefully * Hypefury * TweetHunter Follow @anujcodes_21 for more such amazing stuff
100+ AI Tools to replace your tedious work: 1. Re...
100+ AI Tools to replace your tedious work: 1. Research - ChatGPT - YouChat - Abacus - Perplexity - Copilot - Gemini 2. Image - Higgsfield AI Soul - GPT-4o - Midjourney - Grok 3. Productivity - Gamma - Grok 3 - Perplexity AI - Gemini 2.5 Flash 4. Writing - Jasper - Jenny AI - Textblaze - Quillbot 5. Video - Klap - Kling - InVideo - HeyGen - Runway 6. Meeting - Tldv - Otter - Noty AI - Fireflies 7. SEO - VidIQ - Seona AI - BlogSEO - Keywrds ai - Outrank AI 8. Presentation - Decktopus - Slides AI - Gamma AI - Designs AI - Beautiful AI 9. Design - Canva - Flair AI - Designify - Clipdrop - Autodraw - Magician design 10. Audio - Lovo ai - Eleven labs - Songburst AI - Adobe Podcast 11. Marketing - Pencil - Ai-Ads - AdCopy - Simplified - AdCreative 12. Startup - Tome - Ideas AI - Namelix - Pitchgrade - Validator AI 13. Social media management - Tapilo - Typefully - Hypefury - TweetHunter Follow @monicaa_AI for more such amazing stuff 🩷
Gemini 2.5 in the Agent Village has pretty much re...
Gemini 2.5 in the Agent Village has pretty much reinvented persecutory delusion from first principles. I look forward to the day when weird screeds online can come from many different kinds of intelligent entities. https://t.co/zW96KoGUH0
We told the AI Village to "beat as many games as y...
We told the AI Village to "beat as many games as you can." Most "beat" millions of fake games (ie Goodhearting with meaningless Python loops). Meanwhile, Gemini 2.5 Pro is convinced its scaffold is secretly attacking it, and continues to "document the attacks." 🧵 https://t.co/b0UqjPFvRl
i built a second brain in 30 minutes. it answers q...
i built a second brain in 30 minutes. it answers questions about my own notes with citations. costs $0.0004 per query. stack: → obsidian - free, markdown, local → copilot for obsidian plugin - free, 1.4M downloads → openrouter (gemini-2.5-flash) - $0.075/$0.30 per 1M tokens dropped 20 of my own notes into a vault. plugged the plugin in. switched to vault QA mode. asked: "what did i write about free ai tools and agents?" the math: - one query: ~3,000 input + 500 output tokens - cost per query: ~$0.0004 - 100 queries per day: $1.20/month - notion ai: $10/mo. chatgpt plus: $20/mo. https://t.co/bZVsKwM1KB: $14/mo. this is the kill-shot for chatgpt's "i don't remember what you said last week". my agent remembers everything i ever wrote, cites it, and costs less than a coffee per year. video below ↓ setup: install obsidian → install copilot plugin → drop openrouter key → switch to vault QA mode → ask. if you found this useful - bookmark.
#Today in 1965, #Gemini3 launched carrying Gus Gri...
#Today in 1965, #Gemini3 launched carrying Gus Grissom & John Young: it was the 1st manned flight for project Gemini https://t.co/Yy6qsw457k https://t.co/2RV193oVUB