分类筛选
找到 1605 个 Demo
Google Gemini 3.5 Pro 被曝即将发布,错过了 6 月窗口后消息指向 7 月初...
Google Gemini 3.5 Pro 被曝即将发布,错过了 6 月窗口后消息指向 7 月初 这次升级重点很明确:2M token 上下文窗口、Deep Think 推理模式、编程和 Agent 工作流大幅增强 2M 上下文意味着可以一次性塞进一整个中型代码仓库,Deep Think 则是 Google 对标 o3 的深度推理路线 UI 和前端生成能力也是重点升级方向,Gemini 2.5 Pro 在这块的表现已经让不少开发者转投 Google Claude 4 系列刚稳住局面,GPT-5.6 也在持续迭代,Gemini 3.5 Pro 这个时间点入场压力不小 三家前沿模型的竞争窗口越来越短了,你现在主力用的是哪家?
imagine having gemini 2.5 flash, groq's llama 3.3,...
imagine having gemini 2.5 flash, groq's llama 3.3, cerebras deepseek v4, qwen 3.7, and minimax m3 all behind ONE api endpoint that costs zero dollars 😳 1.7 billion FREE AI tokens every month, one api endpoint, no card needed. there are projects that stack 14+ free ai providers behind a single openai-compatible endpoint with auto-failover when you hit rate limits what they unlock (per month): → 1.7 billion total tokens across all providers → gemini, groq, cerebras, deepseek all in one router → qwen, kimi, minimax, glm, mistral → every major free tier stitched together → when one hits a limit, another takes over → you never see a 429 error again what this replaces: → $200-500/mo in paid api credits → manually juggling 14 api keys → hitting rate limits mid-project → switching endpoints every time a provider changes pricing how to set yours up: step 1: pick a router project > search github for "free-llm-api-router" > or use OmniRoute / 9router > or the openai-compatible router from free-llm-api-resources step 2: configure your providers > add gemini (https://t.co/3DxnIWUfYR) - 1.5m tokens/day > add groq (https://t.co/0Uvpk5uDHT) - fastest llama/qwen inference > add cerebras (https://t.co/GP2py9mDkx) - 1m tokens/day > add deepseek, qwen, minimax, kimi > add nvidia nim (https://t.co/7R7S046Ns3) step 3: set your endpoint > one base url that handles everything > no switching, no key management > plug it into cursor, hermes, claude code, aider, anything step 4: let it fall over > hit gemini's limit? groq takes over > groq busy? cerebras steps in > you keep working, the router handles the rest important: each individual provider has rate limits but the router distributes load. gemini is the backbone (most generous free tier). groq handles speed-critical calls. some routers need self-hosting (docker). the 1.7b figure is the theoretical ceiling across all free tiers combined - real daily usage depends on your workload patterns. your buddy pays $150/mo for one api provider you get 14 providers with auto-failover for $0 bookmark this before the repos get rate-limited too
Google released Nano Banana 2 Lite, a 4-second ima...
Google released Nano Banana 2 Lite, a 4-second image model, alongside Gemini Omni Flash. Image generation usually breaks creative work because every trial costs time, money, and attention. The lighter image model lowers that friction with 4-second outputs at $0.034 per 1K-resolution image. Chaining both models is the real product shape, not either model alone. Nano Banana 2 Lite makes reference images, then Gemini Omni Flash animates them. Google positions it as the replacement for gemini-2.5-flash-image across high-volume developer pipelines. Users still need prompt adherence, stable characters, and readable text during fast visual testing. Gemini Omni Flash extends the workflow from image drafts to editable 10-second video outputs. It accepts text, image, and video inputs, then edits clips through conversation. Pricing: $0.10 per second of video output, matching Veo 3.1 Fast. Gemini Omni Flash currently generates 10-second clips and lacks API audio reference support. Google says the API accepts video references up to 3 seconds, but Gemini Omni Flash does not process them correctly yet.” Interactions API keeps session context, so users can stack 3 sequential edits.
Gemini 2.5 Pro was soo Goated at its time
Gemini 2.5 Pro was soo Goated at its time
used gemini 2.5 pro to build a simple shot counter...
used gemini 2.5 pro to build a simple shot counter for myself + give jordan feedback per shot. https://t.co/GaonBFBGMY
introducing nano banana 2 lite: our fastest, most...
introducing nano banana 2 lite: our fastest, most cost-effective gemini image model yet built for high-velocity developer pipelines, it delivers text-to-image outputs in 4 seconds at just $0.034 per 1K-resolution image swap it into your workflow today via ai studio and the gemini api
Today is the day! We’re officially bringing next-g...
Today is the day! We’re officially bringing next-gen media workflows to the Gemini API and Google AI Studio with a massive dual model launch. Live right now: 🍌 Nano Banana 2 Lite (GA): Our fastest Gemini Image model yet. Built for velocity and high-throughput, it delivers text-to-image outputs in just 4 seconds at $0.034 per 1K resolution image. 🎬 Gemini Omni Flash (Preview): Gemini’s multimodal reasoning meets video. Generate high-quality clips and execute conversational, multi-turn edits with natural language. 🔗 The Combo: Chain them using the Interactions API. Spin up fast reference assets with NB2 Lite, then feed them directly into Omni Flash to animate them. Playgrounds, docs, and SDKs are officially live. Go build! 🛠️
Project Mariner is an early research prototype bui...
Project Mariner is an early research prototype built with Gemini 2.0 that explores the future of human-agent interaction, starting with your browser.
GOOGLE JUST TURNED GEMINI INTO AN AI EMPLOYEE Gem...
GOOGLE JUST TURNED GEMINI INTO AN AI EMPLOYEE Gemini 3.5 Flash can now operate a computer from start to finish—and the business workflows are where this gets serious. What It Can Do: → See your screen and understand what is happening → Click, scroll, type, open tabs, switch apps, and fill forms → Collect information across multiple websites without manual input Business Workflows: ✓ Find warm leads and organize their details in a spreadsheet ✓ Audit competitor content from the last 30 days by engagement ✓ Handle repetitive research, data extraction, and follow-up preparation The important upgrade: Computer Use previously needed a separate standalone model. Now it is built directly into Gemini 3.5 Flash. On the OSWorld benchmark, it scored 78.4—up from 65.1 for Gemini 3 Flash and matching Claude Sonnet 4.6. The best first use is not giving it unlimited control. Give it a predictable task, strict access, a sandbox, and confirmation steps before irreversible actions. AI is no longer just helping you plan the work. It can now sit at the computer and complete it.
google is casually giving developers 1M tokens per...
google is casually giving developers 1M tokens per minute for free 😳 no credit card no subscription just official access through google ai studio what you get for $0: - 1M TPM on gemini 2.5 flash and pro - deep reasoning with pro + ultra-fast inference with flash - native text, image, audio, and video support - instant api key generation in seconds why this is huge: > no fighting strict free-tier limits > no topping up credits just to experiment > no paying middlemen for api access getting started takes less than a minute: 1. go to https://t.co/9vHiKEo4K2 2. sign in with your google account 3. choose flash or pro in the playground 4. generate an api key and start building pro tip: use flash for high-volume workloads and save pro for tasks that need stronger reasoning to get the most out of the free limits the best part? you can access all of this without spending a single dollar free tiers can change anytime, so enjoy it while it lasts bookmark this and grab your api key before everyone else does 👀
This has been very clear for a long time for anybo...
This has been very clear for a long time for anybody working near computer-use. The frontier labs just didn’t care. A short history on this is: until second half of last year only Google out of all big labs had a model capable of using a computer how a human would (understanding xy coordinates of thing it wants to click). OpenAI started catching sometime late last year/early this. Anthropic is still not giving a fuck, at least last time I checked and on public models. Interestingly OpenAI computer use wasn’t too bad, but there was no way of knowing whether they had some specialized model for vision there or used some other representation. Bunch of Chinese labs did specialized vision models that understood interfaces and coordinates. Out of all western labs only Molmo (@natolambert, @finbarrtimbers) and Molmo 2 were usable for this, what’s more is they were VERY good at the task. Before Gemini released the segmentation tool for Gemini 2.5 Flash, it was THE model to use to give a smarter, but blind, model eyes. Building a basic computer-use agent is quite easy (https://t.co/1UL2oa5Wk3, here is a very simple and extendable implementation I did a couple months ago). The main blocker is model coordinate understanding, as most frontier models understand images, but are simply too bad at clicking. General computer-use is simple and will be solved soon if labs want it.
Google just quietly closed the gap between "AI tha...
Google just quietly closed the gap between "AI that talks" and "AI that does." Most people haven't noticed yet. On June 24, Gemini 3.5 Flash got a built-in Computer Use tool no separate model required. → It sees your screen → Clicks, scrolls, types, fills forms → Pulls data across multiple tabs and sites → Loops itself until the task is actually done The number that stands out: 78.4 on OSWorld-Verified (the benchmark for testing real computer tasks). The previous Gemini 3 Flash scored 65.1. That's not an incremental bump. That's a different category of model. ✔ One model that sees, reasons, and acts — no stitching two systems together. ✔ Live now in public preview via the Gemini API and Enterprise Agent Platform. Save this video, you'll want the benchmark numbers next time someone says "AI agents aren't ready yet." Want the SOP? DM me.