릴리스노트

Changelog

총건수

Reve 2.1 is now in the AI Studio (Create) on Abocado AI. It's a 4K image model that landed at #2 overall (Elo 1306) in the Text-to-Image Arena — but the real win isn't the leaderboard, it's the hands-on part: element-level editing, where you change just the headline and leave the product and background pixel-for-pixel, plus the ability to place legible text in several languages right inside the image. Because Reve 2.1 sets an editable layout first and then renders in native 4K rather than pouring the whole picture out at once, both of those come naturally.

What's new

  • Native 4K (16MP) output — resolution that holds up in print and large formats, and stays sharp through repeated edits.

  • Element-level editing — mask one region and re-render it. Change the headline to "40% OFF" while the lighting and product stay untouched.

  • Multilingual in-image text — the long-standing hard problem for image models, now rendered legibly, foreign scripts included.

  • Remix from multiple references — combine different reference images into a new composition.

  • 1:4 to 4:1 aspect ratios, plus auto — from vertical stories to wide banners, with auto picking the shape that fits the request.

Legible text rendered inside the image — Reve 2.1

New models added

  • Reve 2.1 — Text to Image (T2I) — generate 4K images from text, following the prompt literally even in dense scenes.

  • Reve 2.1 Edit — pick one region of an existing image and re-render only that part (element-level editing).

  • Reve 2.1 Remix — compose multiple reference images into a new layout.

Who it's for

It shines when precision matters: posters, ad key visuals, thumbnails, packaging, and UI mockups that need accurate in-image text; dense layouts with many elements; and global campaigns that ship one design across several languages.

Where to try it

Open the AI Studio (Create) on Abocado AI, pick Reve 2.1 among the image models — it carries a 4K badge — and move between text-to-image, edit, and remix in one model. Wrap text you want rendered in quotes, and for edits just mask the region you want to change. Pricing follows your plan — check the app for current credits.

FAQ

What does Reve 2.1 do?

It's a text-to-image and editing model that sets an editable layout first, then renders in native 4K. Accurate multilingual in-image text and element-level editing are its signatures.

How does editing work?

Mask the region you want to change and describe the new content, and only that part re-renders. Everything else stays intact, which makes iterative tweaks — headline, color, material — far easier.

Is the output always 4K?

Yes. Reve 2.1 is a native 4K model, so the default output is always 4K. The auto and 1:4–4:1 aspect ratios set the image's shape (its proportions), not the resolution — whether it's a vertical story or a wide banner, it still renders in native 4K.

Model deep-dive

Why the layout-first approach makes repeated edits and crisp text possible, what the benchmarks mean, and how T2I, edit, and remix work in practice — the full story is in the deep-dive.

The deeper model story

Reve 2.1's layout-first architecture, benchmark breakdown, and the full guide to using it on Abocado AI

Read more

Explore Image Models

Browse all image models on Abocado AI

Explore

Start Creating

Generate and edit 4K images with Reve 2.1 right now

Create

Example image: Reve's official blog (blog.reve.com), cited for this article.

Krea 2 Turbo LoRA is now available on Abocado AI.

New Models

  • Krea 2 Turbo LoRA — Text to Image (T2I) — Apply 1,500 Krea 2 style LoRAs from a gallery with one click, with custom LoRA stacking

Model Details

Krea 2 Turbo LoRA is a text-to-image model that generates images by layering style LoRAs on top of the speed-optimized Krea 2 Turbo base. Abocado AI integrated the entire 1,500-style LoRA collection into a gallery across seven categories — photographic, painterly, illustration, 3D, cinematic, drawing, and graphic — so you pick a thumbnail with your eyes and apply it in one click. The trigger phrase each style needs is appended to your prompt automatically, and styles you like pin to the top of the list in most-recent order. A strength slider and stacking of up to three LoRAs (including ones you trained yourself) are supported. One short line of prompt is all it takes to get the look you want.

Model Deep Dive

Krea 2 Style LoRA features, how-to, and the full 1,500-style collection

Read More

Explore Image Models

Browse all image models on Abocado AI

Explore

Start Creating

Generate images with Krea 2 Turbo LoRA right now

Create

Two Krea 2 Turbo models from the creative AI lab Krea are now live on Abocado AI.

Newly Added Models

  • Krea 2 Turbo (Text to Image) — a 12B diffusion transformer generating native 2K (up to 2048×2048) in about 2 seconds with 8-step inference. 11 aspect ratios × 1M–4M resolution

  • Krea 2 Turbo LoRA (Text to Image) — the variant that runs a custom LoRA on top for generation in your own style. 9 aspect ratios × 1M–2M resolution

Model Details

Krea 2 Turbo is the production checkpoint of Krea 2, the foundation model that independent lab Krea built completely from scratch. The fully post-trained checkpoint — reinforcement learning included — was distilled to compress what usually takes dozens of diffusion steps into just 8. The result is one native-2K image in about 2 seconds. When the wait for each draft disappears, rapid ideation becomes this model's home ground: reshape the prompt, generate several candidates, and choose. It ranks in the global top 10 on the Artificial Analysis text-to-image leaderboard and among the very best from independent labs — released as open weights.

Speed isn't the whole story. It covers a wide, un-"AI-looking" range from film-grain photography to clean studio shots, cinematic stills, and vector illustration, and it renders poster and banner text legibly. When you need a custom style, load your own LoRA on the Krea 2 Turbo LoRA variant and produce with your brand look locked in. Explore fast with Turbo, then finish precision-critical shots with a specialist model — both models are available right now in the Abocado AI Create studio.

Model Deep Dive

Krea 2 Turbo's timeline and key specs, concept and intuition guides, and the full hands-on usage guide

Read More

Explore Image Models

Browse all image models on Abocado AI

Explore

Start Creating

Generate images with Krea 2 Turbo right now

Create

Google DeepMind's three Nano Banana Lite models are now on Abocado AI.

Newly Added Models

  • Nano Banana 2 Lite (Text to Image) — One image in about 4 seconds at the lowest cost tier. The current workhorse, keeping character consistency and in-image text rendering

  • Nano Banana Lite (Text to Image) — The first-generation Lite, for lineage and compatibility

  • Nano Banana Lite Edit (Image to Image) — Conversational targeted editing: tell it to change "just this part" of a reference image

Model Details

Nano Banana 2 Lite isn't a "cheap image generator" — it packs Nano Banana's quality into about 4 seconds at the lowest cost, unlocking high-volume, iterative, near-real-time work. Released by Google DeepMind on June 30, 2026 as the fastest, most cost-efficient Gemini Image model (Gemini 3.1 Flash-Lite Image), it renders an image in roughly four seconds. Speed changes how you work: instead of laboring over a single shot, you pull several options fast, find the direction, and iterate at scale.

Speed didn't cost quality. It keeps the character consistency, legible in-image text rendering, and real-world knowledge that made the original Nano Banana explode in 2025, at 1K resolution across 14 aspect ratios with SynthID and C2PA watermarking. Use the heavier models when a complex composition needs maximum fidelity, and Lite for exploration, volume, and iteration — a pro workflow that splits the job by need. You can also pair it with Abocado AI's Gemini Omni Flash to turn images into video. All three models are available on Abocado AI today.

Model Deep Dive

The three Nano Banana Lite models, plus concept, intuition, usage guides and full benchmark analysis

Read More

Explore Image Models

Browse all image models on Abocado AI

Explore

Start Creating

Generate images with Nano Banana 2 Lite right now

Create

Google DeepMind's Gemini Omni Flash is now on Abocado AI.

Newly Added Models

  • Gemini Omni Flash — Text to Video (T2V) — Generates video with audio from text, grounded in real-world knowledge and physics

  • Gemini Omni Flash — Image to Video (I2V) — Extends a single still image into coherent motion video

  • Gemini Omni Flash — Video to Video (Conversational Edit) — Edits across multiple turns, keeping scene consistency without regenerating

  • Gemini Omni Flash — Reference to Video — Takes text, image, audio, and video references together to guide subject, motion, and style

Model Details

Gemini Omni Flash isn't a "video generator" — it's a unified multimodal model that understands the world and renders video from that understanding. As the first model in the Gemini Omni family unveiled at Google I/O 2026, it processes text, image, audio, and video together in a single forward pass rather than chaining separate models. Switch modes and the output never loses its through-line.

Two things set it apart. First, it generates with knowledge — grounded in Gemini's grasp of history, science, and physics (gravity, fluids), it renders a "protein-folding explainer" accurately, not just plausibly. Second, it lets you edit by conversation — "just this shot at night, keep the rest" — and as instructions stack, the scene remembers what came before, so characters and physics hold. It ranks #1 in human-rated video-editing instruction following and preference. All four modes are available on Abocado AI today.

Model Deep Dive

Gemini Omni Flash's four modes, plus concept, intuition, usage guides and benchmark analysis

Read More

Explore Video Models

Browse all video models on Abocado AI

Explore

Start Creating

Generate video with Gemini Omni Flash right now

Create

ByteDance's Seedance 2.0 Mini is now on Abocado AI.

New Models

  • Seedance 2.0 Mini — Text to Video (T2V) — Generate social video from text prompts, 720P, ~2x faster than Seedance 2.0 Fast, half the cost of Standard

  • Seedance 2.0 Mini — Image to Video (I2V) — Convert images into social video clips, 720P

  • Seedance 2.0 Mini — Reference — Up to 12 combined inputs (images, audio, video) with consistent subject and background fidelity

Model Details

Seedance 2.0 Mini is a lightweight AI video model from ByteDance. It delivers 720P output optimized for Instagram Reels, TikTok, and YouTube Shorts — approximately 2x faster than Seedance 2.0 Fast and at half the cost of Standard, making it practical to produce social content at scale.

Text (T2V), image (I2V), and Reference inputs (up to 6 images + 3 audio + 3 video, max 12 total) are all supported. Aspect ratios: 16:9 · 9:16 · 1:1 · 4:3 · 3:4. Clip length: 4–15 seconds at 24fps. The recommended workflow is to finalize direction with Mini, then generate high-resolution deliverables with Seedance 2.0 Standard when needed.

Full Model Guide

Seedance 2.0 Mini features, workflow guides, and benchmark analysis

Read More

Explore Video Models

Browse all video models on Abocado AI

Explore

Start Creating

Generate social video with Seedance 2.0 Mini right now

Create

Luma Ray 3.2 is now available on Abocado AI. T2V, I2V, V2V, and Reframe — all four modes supported.

Models Added

  • Luma Ray 3.2 Text to Video — generate video from a text prompt

  • Luma Ray 3.2 Image to Video — extend an image into video

  • Luma Ray 3.2 Video to Video — use an existing video as reference to generate a new one

  • Luma Ray 3.2 Reframe — recompose angle, ratio, and perspective of existing video

Launch discount: 20% off all four models.

Where to Use It

Select Ray 3.2 from the V2V, T2V, or I2V tabs in the Abocado AI creation studio.

Model Details

16-keyframe directorial control, 8-person simultaneous performance tracking, ACES2065-1 EXR native output — #1 in 5 benchmarks. Full features, 8 workflow videos, and benchmark analysis.

Model Deep Dive

Ray 3.2 features, 8 production workflow guides, and full benchmark analysis

Read More

Explore Video Models

Browse all video models on Abocado AI

Explore

Start Creating

Generate video with Ray 3.2 right now

Create

xAI's Grok Imagine Video 1.5 is now available on Abocado AI. It currently holds the top spot on the Image-to-Video Arena at Elo 1,473 — ahead of Sora 2, Veo 3.1, Seedance 2.0, and Kling across three independent blind benchmarks. At roughly one-seventh the cost of Sora 2.

What's New

  • Grok Imagine Video 1.5 (I2V) — Animate still images into video. +52 Elo over Video 1.0, the largest single-version improvement in the arena's history. Verified across 3 independent benchmarks.

  • Native audio in a single pass — Sound effects, background music, and lip sync generated simultaneously with video. No post-processing audio step required.

  • Extend from Frame — Select any frame from an existing clip and extend by 6–10 seconds per operation.

  • 4 animation modes (Normal · Fun · Custom · Spicy), 7 aspect ratios.

  • ~25s generation for 6s at 720p (40% faster than Video 1.0).

Why It Matters

The conventional model in AI video has been: top quality costs top dollar. Sora 2 launched at $30/min and people nodded — fair enough. The models that followed kept that expectation in place. Video 1.5 breaks it.

Aurora's autoregressive engine processes frames sequentially — each frame built from the last. Subject position, lighting direction, and camera trajectory stay coherent across the clip. The visual artifacts that make AI video feel "off" often come from frame-to-frame inconsistency; sequential processing reduces those. That architectural efficiency is what delivers Elo +52 and a price point one-seventh of Sora 2 at the same time.

Ceiling: 720p maximum, fixed 24fps. Teams with 1080p delivery requirements should evaluate accordingly. For content creators, marketing teams, and social video production: this is more than enough.

Supported Specs

How to Use

  1. Go to Create → Video Generation.

  2. Select Grok Imagine Video 1.5.

  3. Upload a reference image, enter a motion prompt, and generate.

FAQ

Q: Does it support T2V (text-to-video) as well? Yes. Text-to-video is supported alongside I2V. You can generate video from a text prompt alone, without a reference image — choose whichever fits your workflow.

Q: Can I generate video without the native audio? Yes, audio generation is optional. You can receive video-only output and add audio in post, or let the model generate sound effects, background music, and lip sync in the same pass.

Q: How do I extend a clip that came out too short? Use Extend from Frame. Select the last (or any) frame of the generated clip, and the model extends it by 6–10 seconds. Keep extension passes to a minimum — repeated extensions can accumulate consistency drift over time.


더 깊은 모델 이야기 ▸

Explore video models ▸

Try it now

xAI's Grok Imagine Quality Mode is now available on Abocado AI. Powered by the Aurora MoE (Autoregressive Mixture-of-Experts) architecture, it generates images with maintained text rendering quality at a speed — roughly 4 seconds — where most models stop trying.

Grok Imagine Quality Mode generated example

What's New

  • Grok Imagine Image Quality (T2I) — Text-to-image generation. ~4s latency, up to 10 images per request. Long prompt support (~1,000 characters).

  • Grok Imagine Image Quality Edit (I2I) — Image editing via prompts. Accepts up to 3 reference images in a single request — combine brand assets, mood references, and composition guides at once.

  • 5 technical strengths: volumetric lighting (God Rays, subsurface scattering), material fidelity, multilingual in-image text, complex spatial composition, brand and cultural element accuracy.

Why It's Different

Diffusion models treat text prompts as external guidance — a direction to approximate, not a specification to follow. That's the structural reason fast diffusion models struggle with legible in-image text: the letters aren't part of the generation process, they're a condition on top of it.

Aurora MoE is autoregressive. Text tokens and image tokens are processed in the same pipeline. The words in your prompt — including the ones you want rendered visually — are part of how the image is built, not instructions layered on top. The result is that banners, posters, and ad creatives with legible, well-placed text come out in roughly 4 seconds.

This is not the best text-rendering model available. ChatGPT Image Gen and comparable dedicated tools are still ahead on absolute accuracy. But at this speed, this level of text fidelity is rare. If your workflow involves text-heavy images at volume, the tradeoff between speed and quality collapses.

Supported Specs

How to Use

  1. Go to Create.

  2. Select Grok Imagine Image Quality (T2I) or Grok Imagine Image Quality Edit (I2I).

  3. Enter your prompt and generate.

For I2I: upload up to 3 reference images. Combine brand colors, mood references, and layout guides in one request.

FAQ

Q: What's the difference between Speed Mode and Quality Mode? Speed Mode prioritizes lower latency for fast iteration. Quality Mode focuses on image fidelity — volumetric lighting, material textures, and text rendering. If you're generating final-quality banners or posters, Quality Mode is the right choice.

Q: Is Quality Mode the best model for in-image text? Not absolutely. Dedicated text-rendering models like ChatGPT Image Gen still outperform it on raw text accuracy. Quality Mode's advantage is the combination: it holds a high enough standard for text while delivering results in ~4 seconds — a tradeoff most text-strong models can't match on speed.

Q: How should I use the 3 reference images in I2I? Assign each image a different role. Example: 1st → brand color palette, 2nd → mood/atmosphere reference, 3rd → composition or layout guide. This lets you communicate visual intent that's hard to express in words.


더 깊은 모델 이야기 ▸

Explore image models ▸

Try it now

Ideogram 4.0, Ideogram's first open-weight image model, is now available on Abocado AI. Plenty of models make good-looking photos — this one's day job is different: rendering the text, layout, and brand colors inside an image exactly as you specify. In real design-work evaluations, nearly half of designers (47.9%) ranked it their first pick.

What's New

  • Ideogram 4.0 (text-to-image) — a 9.3B-parameter open-weight model. Within days of its June 3, 2026 release it ranked #1 among open-weight image models and #2 overall (only GPT and Gemini above it).

  • In-image text accuracy (OCR) of 0.97 — multi-line, multi-font type rendered precisely at native 2K.

  • Colors lock to up to 16 hex codes per image and element positions to coordinates — the same instruction returns the same result, every time.

Why It's Different

To most models, a prompt is a request — a similar blue, roughly at the top, different every run. Ideogram 4.0 was trained on structured specifications, so where the text sits and what color it wears is honored like a contract. The gap shows most in finished designs with type in them — posters, thumbnails, ads, logos — and especially brand or campaign series where layout and color must stay locked.

How to Use

  1. Open Create (/ai/create) or Explore → Image (/ai/explore/image).

  2. Pick Ideogram 4.0 from the model selector.

  3. Describe the design and the exact copy you want in it — the more specific your text and layout instructions, the more this model shines.

The Full Story — One Deep Read

In an era when Nano Banana and ChatGPT image write text just fine, why did designers pick this one? The difference between "it can write" and "it does what you specify" — unpacked in a single story.

FAQ

  • How is this different from other image models? It's not whether it can write text, but whether the text lands at your specified position, color, and font — identically every time. That matters most in consistency-driven design work.

  • Is it open source? Half — the code is free (Apache-2.0), the weights are non-commercial. On Abocado AI you use it with none of that overhead.

  • What is it best for? Finished designs with type: posters, thumbnails, ads, logos — and campaign series with locked brand colors and layouts.

Related

NVIDIA Cosmos 3 Super — text-to-image from NVIDIA's newly released omnimodal world model — is now available on Abocado AI (Abocado AI), one of the first platforms to offer it. Unlike a model that just paints pretty pictures, a world model learns how the world actually works — light, weight, cause and effect — and that understanding shows up in what it generates.

What's new

  • Cosmos 3 Super (text-to-image) is part of NVIDIA's Cosmos 3 family — a single framework that unifies language, image, video, audio, and action, built for "Physical AI."

  • Released in June 2026 as an open frontier foundation model, Cosmos 3 was ranked the best open-source text-to-image and image-to-video model by Artificial Analysis at launch.

  • On Abocado AI it arrives fast — generate with it the moment it's news, before everyone else is even searching for what a "world model" is.

Why a 'world model' is different

Most image models map text straight to pixels. A world model reasons about scene physics first — how light falls, how objects sit in space — and then renders. That's the shift from "AI that draws" to "AI that understands the world," and it's why the same Cosmos 3 family also drives robots and simulators, not just images. You feel it as outputs that simply hold together.

How to use it

  1. Open Create (/ai/create) or browse Explore → Image (/ai/explore/image).

  2. Choose NVIDIA Cosmos 3 Super in the model selector.

  3. Write your prompt — prompt expansion is on by default, so a short idea gets enriched automatically. Add a negative prompt to steer away from what you don't want.

  4. Generate up to 4 images per run (PNG or JPEG).

Go deeper — a 3-part series

We unpacked why this model is called "the future of generative AI" in three short reads — intro, concept, and intuition. Read them in order and the bigger picture clicks:

FAQ

  • Is this just another text-to-image model? No — it's a world model that reasons about physics, brought here for image generation.

  • Do I need to write long prompts? No — prompt expansion is on by default and fills in detail from a short prompt.

  • What output do I get? Up to 4 images per run, in PNG or JPEG.

Related

Microsoft's first image models — MAI-Image-2.5 (generate) and MAI-Image-2.5-edit (edit) — are now available on Abocado AI (Abocado AI). Fresh from launching at No. 2 for image editing and No. 3 for text-to-image on LMArena, both are live in the image apps and ready to use — no post-processing required.

What's new

  • MAI-Image-2.5 (text-to-image) turns a prompt into a detailed, coherent image with notably strong text rendering, product imagery, and prompt adherence. It reasons about scene structure, lighting, scale, and perspective, so results hold together without retouching.

  • MAI-Image-2.5-edit (image editing) makes precise, localized changes — swap an object, fix text, or remove motion blur in just the region you point to, while the rest of the image stays exactly as it was.

  • A +75-point overall jump over MAI-Image-2, with the biggest gains in text rendering (+107) and cartoon, anime & fantasy (+90).

Why it matters

Image AI was Google's (Nano Banana) and OpenAI's (GPT-Image) stage — and Microsoft landed in the top three on its first try, beating Nano Banana 2.1 in blind editing tests across 12 categories. For you that means cleaner text in posters, edits that respect the original lighting and shadows, and a face that stays the same person across changes in pose, expression, and viewpoint.

How to use it

  1. Open Create (/ai/create) or browse Explore → Image (/ai/explore/image).

  2. In the model selector, choose MAI-Image-2.5 to generate, or MAI-Image-2.5-edit to edit an existing image.

  3. For edits, upload your image and describe the change for just the area you want — the rest is preserved.

  4. Generate — pick an aspect ratio (up to 21:9 wide or 9:16 tall) and up to 4 variations per run.

Supported & limits

  • Output formats: PNG, JPEG, WebP. Up to 4 images per generation; the editor also accepts reference images for guided changes.

  • Availability can vary by plan — check the model in Create or your pricing page.

FAQ

  • What's the difference between the two models? MAI-Image-2.5 creates images from text; MAI-Image-2.5-edit changes an existing image in a targeted region.

  • Will editing change the whole picture? No — edits are localized to the area you specify; the rest stays untouched.

  • Does the same person stay consistent across edits? Yes — faces hold across pose, expression, and viewpoint.

Related