Category Archives: AI/ML

GPT Astra driving AE & Premiere Pro

Desire to know more intensifies…

“Everyone in the world now has a 3D designer at their fingertips”

Insane that 3D worlds can be conjured straight from language, then rendered natively or used as guidance for AI-based rendering. Check out this demo from my former Adobe teammate Tom:

Days of Miracles & Wonder, man—always.

Behind-the-scenes thread:

And yet more magic:

Photoshop combines generative imaging & masking

It seems I’m on something of a Photoshop kick lately (something something you can take the boy out of West 10, but…), so I’ll mention a couple of interesting-sounding enhancements that have landed in the latest release:

  • Instruct Edit with Masks, powered by Firefly Image 5, understands the full context of your image so you no longer need to precisely mark every edit area yourself. Just describe the precise edits you’re looking for — such as opening closed eyes or placing a hat on a specific person’s head — with simple prompts. The unmasked areas remain untouched so your other important elements, like faces, logos, or brand assets, are protected.
  • Light Adjustment Layer gives you professional-grade lighting controls — Exposure, Contrast, Highlights, Shadows, Whites, and Blacks — in a new non-destructive adjustment layer, so you get Camera Raw-level results without ever leaving your Photoshop layers workflow.
  • Markup lets you visually communicate edits by drawing directly on an image to show the model what you want — select areas to recolor, sketch arrows to indicate position, or brush in rough shapes to suggest new elements — so you can get results that more closely match your vision while reducing the ambiguity of text-only prompts.

Omni Flash: Poolside AI

See, I just take Gemini to the dog park, but my teammate Genevieve took it all the way to the pool on vacation. There are levels to this stuff. 🙂

But seriously: this is a really nice visualization of the new clip-extension feature, among other things. That’s how the scene can run a full 30 seconds.

Higgsfield relighting looks wild

I’ll say it again: Forget the video component per se: interactivity like this is either the future of Photoshop, or it’s the replacement for Photoshop. There’s no third way.

Gemini Omni 1.1 Flash is here! 4k, clip extension and more

Come build beautiful things with higher resolutions (and lower: 360p is great for fast, cheap drafts), clip extension (up to 10s at a time, up to 40s total), better support for audio and video references, improved frame interpolation (specifying first/last frames), and more. Let’s tell great stories together. 🙂

From the official blog:

  • More creative control with start and end frames. This helps keep the characters and narrative consistent as you thoughtfully transition between frames.
  • Elevate your rough cuts into polished pieces. Export crisp video in 1080p or 4K, ready for high-end digital, social, or broadcast editing workflows.
  • Draft videos quickly and upscale for quality. Test out concepts and compositions at a faster, lower-credit 360p resolution before committing to a full-resolution render. Once you have a clip you’re happy with, you can then download it in 720p resolution. This is particularly helpful in the Google Flow app — you can quickly draft videos on your phone on the go, and upscale your favorite versions.

See this thread for great examples, and please let me know if you have questions or requests. Onward!

Fun with Omni peech

(plural of pooch: “peech”; same for spouse/speece)

No doods were frightened in the making of this fiery display, made with our Google Omni Flash model:

Nor were any steamed:

Amazing realtime relighting

Wow—provided quality & resolution can be made high enough, tell me how soon tech like this can come to Photoshop & Lightroom:

Getting buff in Yellowstone

Given that we’re taking our eldest son to start life as a University of Colorado Buffalo tomorrow (!), it was only fitting that we communed with a few real buffs in Yellowstone over the last few days. Here’s an Insta gallery:

 

 
 
 
 
 
View this post on Instagram
 
 
 
 
 
 
 
 
 
 
 

 

 

A post shared by John Nack (@jnack)

Bonus wildlife from the trip: in Boise we dropped by the The World Center for Birds of Prey and got to see this handsome crew in action. (Inside-baseball detail: hopefully you’d never know that I was obliged to photograph the first owl shown through a dense set of bars on his enclosure. Nano Banana in Photoshop to the rescue!)

Some amazing recent animations

It’s a very imperfect analogy, but I feel like as AI video fully passes the “can this pass for real” test, we’ll break into new & more interesting territory—much as painting did once photography took the “does this replicate reality” crown.

Here’s a handful of fun, beautiful animations rendered in a variety of styles. The fact of them being powered by new models is kind of incidental—as it should be: all that matters is what moves people & what helps artists do that.

Taking Nano Banana on the road

Greetings from the midst of our bittersweet (but mostly very sweet!) roadtrip to take Finn off to college in Boulder. Me being me, I of course brought our Lego selves to use in making arguably cringey (#jeezdad) family pics. I don’t have a precise match for our newly modded van, so I brought the closest equivalent & then gave Gemini a reference image to use with Nano Banana. Not bad, robot—not bad at all:

Head-spinning photo->3D

I have no words for this kind of witchcraft. In case the embedded tweet doesn’t show up in English, here’s a translation:

Wow. Maciej Dobrodziej – a digital creator – has “brought to life” one of Warsaw’s most famous photographs. The photo of a girl running in the rain on Puławska Street was taken by Zbigniew Siemaszko in 1968. Fun fact: After many years, it was possible to find Ms. Grażyna, who is the heroine of the photo. She recognized herself upon seeing… the photograph on FB.

Elsewhere, Bytedance helps you explore similarly, specifying camerawork just by drawing lines on an image:

Wait for me, Penelope…

“A.I. sing of arms and the man…” Wait—that was the Aeneid, but Imma go for it.

This reimagining of The Odyssey—nailing everything from casting to an apparently generated period-accurate song that actually slaps—demonstrates my long-running contention that when things become amazing enough, we can’t even process them.

People see this and neither marvel at the incredible state of technology (which has leapt forward in just the last couple of weeks), nor point out little shortcomings here & there, nor even get into another fruitless battle about the ethics of AI. Instead, at least from where I sit, I see them simply commenting on the concept & storytelling.

Which, I think, is how it should & will be.

Put your hands together—literally—for the music of MediaPipe

It’s so fun to see work from a past life enabling whole new modes of expression!

Taking a different approach to hand tracking & creativity, check out this Omni experiment:

From Nano Banana to Beast Mode

As I walked the dogs past our new camper van a few weeks back, Pinterest happened to send me some vintage World War II aircraft art. I thought, heh, how nuts would it be to give the van shark teeth? I snapped a quick pic of the van, popped it plus one of the shark mouth designs into the Gemini app, and had Nano Banana mock up the possibilities:

This helped me bring the inimitable & indulgent Margot onboard with the idea, and soon enough I found myself in Illustrator, recreating the art the old fashioned, point-by-point way. Upon seeing the design, the local graphics shop advised me on some needed mods, so I hopped into Photoshop to oblige, invoking a dash of Generative Fill. Check out the (very real!) results.

To me this is just how AI ought to work—not as a creative replacement, but rather as an accelerant that lets us try more things & communicate ideas better.

Flow to the Upside Down

Fill your head with sweet 80’s synth chords, using Google Omni Flash to reimagine suburbia as a sci-fi dreamscape:

Joy-scrolling 3D

I love the art direction on this site, where the pace of animation is controlled by your scrolling:

Meanwhile the “Scroll World” Claude skill promises to interview you (!), then whip something up with the help of generative text-to-video tech (gotta get Omni in there!).

 

How it works is intriguing enough to quote at length from the GitHub page:

—–

It leans on Higgsfield for the art: cohesive isometric diorama scenes (GPT Image 2 — via Higgsfield, or the Codex CLI on a ChatGPT subscription) and the camera flights themselves (Seedance or Kling image-to-video — only models that can frame-lock a seam), scrubbed by scroll position — the same technique behind Apple’s scroll-through product pages. The camera genuinely moves; scroll only drives time. It’s framework-agnostic: you get the Higgsfield pipeline, the prompt templates, and a portable vanilla-JS scrub engine that drops into plain HTML, Next.js, Vue, or a Python-served page — nothing assumes a stack.

When invoked, the skill:

  1. Interviews you — the subject/industry + pitch, a brand kit (import from a URL, hand it over, or have it proposed), art direction, the ordered scenes the camera visits, whether you want the mobile version (a second chain rendered natively in 9:16 portrait — composed for phones, not a crop of the landscape film), and the budget — render tiers and stills source shown with estimated credit costs, approved before anything generates.
  2. Generates the assets — one still per scene, one “dive-in” camera clip per scene, and the connector clips that join consecutive scenes, generated from the actual rendered frames of their neighbours so every seam is frame-identical. Mobile opt-in renders a parallel portrait chain the same way, frame-locked against its own 9:16 renders.
  3. Wires it up — a config-driven scroll engine that plays the whole chain as one flight, serving the portrait clips and posters automatically on phones.

Quick tutorial: Create with Omni inside Adobe Boards

Companies are imaginary. Creativity is real.

I love seeing my old friend & Adobe design partner Dave (mentioned innumerable times here over the years) showing off the expressive power of the collaboration so many Adobe & Google friends have been nurturing over the last several months. We have the first fruits showing now in Firefly Boards (try it now!), and there’s so much more we have in mind.

I’ve never cared much about the company name in my email address. All that stuff is always just a means to an end—a way of aligning incentives so we can help people make the world more fun and beautiful. I feel so grateful to have found a spot to stand among friends & continue this work—and we’re just getting started.

“All of this has happened before, and all of this will happen again”

…So says my friend Chris Perry, who welcomed me to Google 12 years ago and who was essential in shipping the first feature I worked on there—face painting for the 2014 World Cup:

Smash cut to 2026: Chris has just left Big G to start his own company, but we’re still inviting people to paint their faces—and to do so much more—for the World Cup. Opening Gemini today, I saw this smorgasbord of Nano Banana-powered templates:

Finn & Henry (pictured up top) are now far too old and, critically, too cool to abide my applying any patriotic AI to them—but you can give it a try with pics of you & yours. 🙂

Remix your photos using Gemini Omni, right inside Google Photos

The fun has begun:

Located in the Create tab — your central hub for creativity in Google Photos — Video Remix helps you make inspired content in seconds. Instantly apply cinematic relighting to spruce up a dark clip, swap out a plain background for something fun, or add artistic treatments, such as watercolor, raw sketchbook and oil painting effects.

Video Remix starts rolling out today to eligible Google AI Plus, Pro and Ultra subscribers in select countries.

Use Google video models right within Premiere Pro!

I love seeing a plan come together. 😀 Check out this new Premiere Pro integration with Google Veo:

The Google Omni Flash model isn’t there yet, but we’ll work on connecting the dots. As Mark Twain might’ve said, “Some Adobe menus are so long, they have perspective,” and I’d love to get back to contributing to that phenomenon. 😉

In the meantime, my old partner Dave Werner is cooking with Omni inside Firefly Boards—where you can try it yourself:

[Via Bill Hensler]

Fun with Omni: Changing cams, plus Wolverine

Karen X. Cheng shows a subtle but powerful model capability:

 
 
 
 
 
View this post on Instagram
 
 
 
 
 
 
 
 
 
 
 

A post shared by Karen X (@karenxcheng)

Meanwhile Christian Cantrell is looking sharp, is not downright superheroic:

Come build with Gemini Omni Flash!

I’m thrilled to say that the first big launch of my Chapter 2 at Google is here! You can now build on Gemini Omni Flash in Google Enterprise Agent Platform (aka Vertex); see docs.

TBH I was so busy helping get the release out the door, and then taking some much needed rest over the Fourth of July break, that I’ve hardly had a chance to post useful info. I’ll fix that soon! In the meantime, here’s our little intro sizzle reel:

Of beat labs & photo shoots

It’s always cool to see how creators are embracing new tools:

Quick tour: Creating Flow tools with natural language

My teammate Anika & I got to meet the other day with a really big creative brand the other day (more details to share soon, I hope), and they got excited about delivering super focused, relevant experiences for their customers by building on the Flow agent & apps. Here’s Anika offering a concise tour of how to create, share, and remix the latter:

“Google Just Turned Street View Into a Video Game”

As Bilawal puts it,

At Google I/O 2026, DeepMind shipped Maps Imagery Grounding for Genie 3 — their real-time world model can now generate interactive 3D worlds conditioned to any of the 280 billion Street View images Google has captured over 20 years. Pick a location on Google Maps, choose a style, drop in a character, and walk around.

Check out his accessible & illuminating tour of the new tech:

Omni Teapot

My 16yo is lowkey impressed that at Adobe I got to work with Utah Teapot creator Martin Newell. At this point, anything that impresses a teen is very welcome. 🙂

I wonder what he’d think of Gemini Omni turning real teapots into geometry just by saying the word:

Puppetry + AI FTW: Behind the Scenes with Timmy TPU

I love the blend old-school puppetry, 3D animation, Gemini Omni, and the latest experimental video tools that went into creating TPU Training Day, the short film that debuted during Google I/O 2026.

I know you’ve heard it a million times, but it bears repeating: AI isn’t a substitute for human creativity, or in many cases even for traditional techniques. It’s just a whole new toolbox that can multiply our expressive powers.

And here’s the film itself:

Check out Google Flow Agent

Did I have “Google makes cool, extensible, AI-powered creative tools” on my 2026 Bingo card? I did not—and I’m happy to be wrong! Check this out:

According to the docs, you can use the Agent to:

  • Brainstorm and plan: Chat with the Agent to outline storyboards, develop visual mood boards, and turn high-level concepts into actionable prompts.
  • Generate new media: Ask the Agent to generate videos or images and select the best model to generate with.
  • Edit assets directly: Ask the Agent to edit selected media from your project.
  • Batch generate: Ask the Agent to create multiple variations of an asset at once.
  • Organize your assets: Ask the Agent to rename specific files, group selected media into a new Collection, or archive unused assets.
  • Add context & references: Drag media into the Agent prompt box from your device or project. You can also select multiple assets and let the agent know which media you are referring to.

3D typography using Omni + Flow

Check out this cool little technique:

This is especially wild when you consider where typography stood just a couple of years ago—for which I’ll forever be kinda nostalgic. 🙂

A beautiful moment of expressivity unlocked

Despite—or maybe because—of my line of work, I have some genuinely mixed emotions about AI. Is it about empowerment, devaluation, theft, magic? Yes. It is all, as my wife would say of me, A Lot™.

Alongside whatever else it may be, however, the tech can be a genuine enabler of human expressivity. If you don’t believe me, just take 90 seconds to read & watch this heartfelt moment:

Vibe-code your own VFX apps & more Google Flow Tools

Democratize all the apps!!

I think this new platform will be a major sleeper hit:

With Google Flow Tools, you can build creative workflows customized to fit your creative process. Explore a gallery of premade Tools built by creatives, remix existing ones to fit your needs, or create your own from scratch by just typing a description of what you want to create. You can shape and iterate on these Tools fluidly, adjusting them for individual projects, or singular clips and images. All users can explore Tools in Google Flow, and Google AI subscribers can create custom Tools from scratch or remix existing ones.

Check out the ways some artists have been spinning up tools & putting them to work:

Google Earth + Omni = Drone magic

My friend Bilawal, who used to work in Google’s Geo group (Earth, Maps, and more), has created an eye-popping faux-drone video using Omni Flash:

Here’s another exploration, inspired by Bilawal’s:

Awesome examples of Omni video transformation

This is such a wild, game-changing feature:

I think Carlos gets it exactly right: “I think many are focusing on the wrong aspect of the Gemini Omni model when comparing it to Seedance 2.0, since conceptually they are entirely different things. This is a model for editing videos (like Nano Banana) like we’ve never had before!