Matan Cohen-Grumi (see previous) asks, “What if music icons walked into their art?”
And on a much more ridiculous tip, there’s the Belt Squared & beyond!
— Awful Taste But Great Execution (@AwfulButGreat) January 4, 2025
Matan Cohen-Grumi (see previous) asks, “What if music icons walked into their art?”
And on a much more ridiculous tip, there’s the Belt Squared & beyond!
— Awful Taste But Great Execution (@AwfulButGreat) January 4, 2025
Check out the latest from Topaz:
Topaz really cooked with their new upscaling model called “redefine” — basically every CSI “enhance” meme you’ve seen IRL.
Settings:
– 4x Upscale
– Creativity: 2
– Texture: 3
– No promptIt’s basically the Topaz take on the magnific style of “creative upscaling” where you use… pic.twitter.com/T7dLoAjFJt
— Bilawal Sidhu (@bilawalsidhu) December 17, 2024
Alternately, you can run InvSR via Gradio:
Image super-resolution model just dropped! Superior results even with a single sampling step.
InvSR: Arbitrary-steps Image Super-resolution via Diffusion Inversion. pic.twitter.com/gS7uoGwnQ8
— Gradio (@Gradio) December 16, 2024
I’ve long wanted—and advocated for building—this kind of flexible, spatial way to compose & blend among ideas. Here’s to new ideas for using new tools.
Supporting Non-Linear Exploration
Creative exploration rarely follows a straight line. The graph structure naturally affords exploration by allowing users to diverge at various points, creating new forks of possible alternatives. As more exploration occurs, the graph grows… pic.twitter.com/Yq18Caj94T
— Runway (@runwayml) December 2, 2024
It’s a touch odd to me that Meta is investing here while also shutting down the Meta Spark AR lens platform, but I guess interest in lenses has broadly faded, and AI interpretation of images may prove to be more accessible & scalable. (I wonder what’ll be its Dancing Hot Dog moment.)
I love seeing exactly how Chad Nelson was able to construct a Little Big Planet-inspired game world through some creative prompting & tweening in Open AI’s new Sora video creation model. Check out his exploratory process:
View this post on Instagram
Director Matan Cohen-Grumi shows off the radical acceleration in VFX-heavy storytelling that’s possible through emerging tools—including Pika’s new Scene Ingredients:
For 10 years, I directed TV commercials, where storytelling was intuitive—casting characters, choosing locations, and directing scenes effortlessly. When I shifted to AI over a year ago, the process felt clunky—hacking together solutions, spending hours generating images, and… pic.twitter.com/pJUamLFgWI
— Matan Cohen-Grumi (@MatanCohenGrumi) December 18, 2024
Check out this fun little toy:
Instead of generating images with long, detailed text prompts, Whisk lets you prompt with images. Simply drag in images, and start creating.
Whisk lets you input images for the subject, one for the scene and another image for the style. Then, you can remix them to create something uniquely your own, from a digital plushie to an enamel pin or sticker.
Meet Whisk! Our new experiment that lets you use images as prompts to visualize your ideas and tell your story. Try it now: https://t.co/BR1z7gmDs6 pic.twitter.com/2zrPLQZlga
— labs.google (@labsdotgoogle) December 16, 2024
The blog post gives a bit more of a peek behind the scenes & sets some expectations:
Since Whisk extracts only a few key characteristics from your image, it might generate images that differ from your expectations. For example, the generated subject might have a different height, weight, hairstyle or skin tone. We understand these features may be crucial for your project and Whisk may miss the mark, so we let you view and edit the underlying prompts at any time.
In our early testing with artists and creatives, people have been describing Whisk as a new type of creative tool — not a traditional image editor. We built it for rapid visual exploration, not pixel-perfect edits. It’s about exploring ideas in new and creative ways, allowing you to work through dozens of options and download the ones you love.
And yes, uploading a 19th-century dog illustration to generate a plushie dancing an Irish jig is definitely the most JNack way to squander precious work time do vital market research. 🙂

I’m a near-daily user of Ideogram to create all manner of images—mainly goofy dad jokes to (ostensibly) entertain my family. Now they’re enabling batch creation to facilitate creation of lots of variations (e.g. versions of a logo):
Check out this wild video-to-video demo from Nathan Shipley:
Sora Remix test: Scissors to crane
Prompt was “Close up of a curious crane bird looking around a beautiful nature scene by a pond. The birds head pops into the shot and then out.” pic.twitter.com/CvAkdkmFBQ
— Nathan Shipley (@CitizenPlain) December 10, 2024
Just a taste of the torrent the blows past daily on The Former Bird App:
This might be the world’s lowest-key demo of what promises to be truly game-changing technology!
I’ve tried a number of other attempts at unlocking this capability (e.g. Meta.ai (see previous), Playground.com, and what Adobe sneak-peeked at the Firefly launch in early 2023), but so far I’ve found them all more unpredictable & frustrating than useful. Could Gemini now have turned the corner? Only hands-on testing (not yet broadly available) will tell!
Nice to see this progress. (FWIW Microsoft Designer features similar tech; just putting that out there. :-))
New in the @Photoshop Beta! The Object Selection tool just got supercharged! pic.twitter.com/VrfxCQa84W
— Howard Pinsky (@Pinsky) December 9, 2024
Karen X, back doing crafty Karen X things:
AI painting tutorial
Edited on my Intel AI PC – the DELL XPS 13 powered by Intel Core Ultra #ad pic.twitter.com/wcqpR3RhFk
— Karen X. Cheng (@karenxcheng) December 10, 2024
Diffusion models are ushering in what feels like a golden(-hour) age in relighting (see previous). Among the latest offerings is LumiNet:
[6/7] Here are a few more random relighting!
How accurate are these results? That’s very hard to answer at the moment But our tests on the MIT dataset, our user study, plus qualitative results all point to us being on the right track.
It’s like we’ve cracked open a… pic.twitter.com/1FNlz8S9Fk
— Anand Bhattad (@anand_bhattad) December 5, 2024
Zevia cleverly mocks Coke’s use of AI to generate its recent commercial:
Looks like Body Armor had a similar idea a few months back:
What if your design tool could understand the meaning & importance of words, then help you style them accordingly?
I’m delighted to say that for what I believe is the first time ever, that’s now possible. For the last 40 years of design software, apps have of course provided all kinds of fonts, styles, and tools for manual typesetting. What they’ve lacked is an understanding of what words actually mean, and consequently of how they should be styled in order to map visual emphasis to semantic importance.
In Microsoft Designer, you can now create a new text object, then apply hierarchical styling (primary, secondary, tertiary) based on AI analysis of word importance:


I’d love to hear what you think. You can go to designer.microsoft.com, create a new document, and add some text. Note: The feature hasn’t yet been rolled out to 100% of users, so it may not yet be available to you—but even in that case it’d be great to hear your thoughts on Designer in general.
This feature came about in response to noticing that text-to-image models are not only learning to spell well (check out some examples I’ve gathered on Pinterest), but can also set text with varied size, position, and styling that’s appropriate to the importance of each word. Check out some of my Ideogram creations (which you can click on & remix using the included prompts):

These results of course incredible (imagine seeing any of this even three years ago!), but they’re just flat images, not editable text. Our new feature, by contrast, leverages semantic understanding and applies it to normal text objects.
What we’ve shipped now is just the absolute tip of the iceberg: to start we’re simply applying preset values based on word hierarchy, but you can readily imagine richer layouts, smart adaptive styling, and much more. Stay tuned—and let us know what you’d like to see!
I’m still catching up from Thanksgiving, obvs; enjoy these tasty leftovers (all links to demo vids on Twitter—yes, always “Twitter”):
Speaking of Kling, the new Motion Brush feature enables smart selection, generative fill, and animation all in one go. Check out this example, and click into the thread for more:
Kling AI 1.5 Motion Brush is incredible.
You can give different motions to multiple subjects in the same scene.
Game changing control and quality
6 wild examples: pic.twitter.com/sEDbNC1iPq
— Min Choi (@minchoi) November 30, 2024
Accurately rendering clothing on humans, and especially estimating their dimensions to enable proper fit (and thus reduce costly returns), has remained a seductive yet stubbornly difficult problem. I’ve written previously about challenges I observed at Google, plus possible steps forward.
Now Kling is promising to use generative video to pair real people & real outfits for convincing visualization (but not fit estimation). Check it out:
Kling AI just dropped AI Try-On.
Now anyone can change outfits on anyone.
8 wild examples:pic.twitter.com/EKoYjKxTRd
— Min Choi (@minchoi) November 30, 2024
The product provides pre-set models and clothing.
But you can also upload your own – making anyone model anything.
Here’s me in a tank top and jeans I found online pic.twitter.com/FYX3QvHxP0
— Justine Moore (@venturetwins) November 29, 2024
We present FlipSketch, a system that brings back the magic of flip-book animation — just draw your idea and describe how you want it to move! …
Unlike constrained vector animations, our raster frames support dynamic sketch transformations, capturing the expressive freedom of traditional animation. The result is an intuitive system that makes sketch animation as simple as doodling and describing, while maintaining the artistic essence of hand-drawn animation.
Oh, I love this one!
FlipSketch can generate sketch animations from static drawings using text prompts!
Links ⬇️ pic.twitter.com/1XPzkWfaEl
— Dreaming Tulpa (@dreamingtulpa) November 22, 2024
I’m finding the app (which is free to try for a couple of moves, but which quickly runs out of credits) to be pretty wacky, as it continuously regenerates elements & thus struggles with identity preservation. The hero vid looks cool, though:
BlendBox AI: Seamlessly Blend Multiple Images with Ease
It makes blending images effortless and precise.
The real-time previews let us fine tune edits instantly, and we can generate images with AI or import our own Images.
Here is how to use it: pic.twitter.com/9LyVF8x8qN
— el.cine (@EHuanglu) November 19, 2024
Check out LLaMA-Mesh (demo):
Nvidia presents LLaMA-Mesh
Unifying 3D Mesh Generation with Language Models pic.twitter.com/g8TTaXILMe
— AK (@_akhaliq) November 15, 2024
Ignoring the misguided (IMHO) contents of the surrounding tweet, I found these four minutes of commentary to be extremely sharp & well informed:
I wonder whether such statements are psychological defense mechanisms such as repression and denial.
In any case, some people will very soon realize that reality is different from their illusory wishful thinking. pic.twitter.com/Y9mDkAZToI
— Chubby♨️ (@kimmonismus) November 15, 2024
Creative control to the people! I can’t wait to try this out:
the wait is over, our new AI trainer is out!
it comes with upgraded quality and hundreds of community styles you can use in your generations.
full tutorial below pic.twitter.com/oFtaiEmCS9
— KREA AI (@krea_ai) November 14, 2024
Man, I miss working with these guys & gals…
We present ReCapture, a method for generating new videos with novel camera trajectories from a single user-provided video. Our method allows us to re-generate the source video, with all its existing scene motion, from vastly different angles and with cinematic camera motion.
They note that ReCapture is substantially different from other work. Existing methods can control camera either on images or on generated videos and not arbitrary user-provided videos. Check it out:
Paul Trillo relentlessly redefines what’s possible in VFX—in this case scanning his back yard to tour a magical tiny world:
Getting my hands dirty with 30 Gaussian splats scanned in my garden. Is this the most splats ever in a single shot?
Made with the support of @Lenovo @Snapdragon and the new Gaussian Splatting plugin by @irrealix pic.twitter.com/ezXo6MMnQi
— Paul Trillo (@paultrillo) October 3, 2024
Here he gives a peek behind the scenes:
How I created the love letter to the garden and bashed together 30 different Gaussian splats into a single scene pic.twitter.com/OKxDFtK8uE
— Paul Trillo (@paultrillo) October 18, 2024
And here’s the After Effects plugin he used:
DifFRelight can change flat-lit facial captures into high-quality images and dynamic sequences with complex lighting!
It uses a diffusion-based model for precise lighting control, accurately showing effects like eye reflections and skin texture.https://t.co/iXZjpVOxFx pic.twitter.com/Rwj7jTBnRd
— Dreaming Tulpa (@dreamingtulpa) October 23, 2024
Check out this impressive use of the new “retexture” feature, which enables image-to-image transformations:
Wow!
Midjourney’s edit and retexture features are incredible!
I retextured my profile image using some of my sref codes and animated it with LumaLabs.
The new images are stunning, and the animated video looks even better.
You can apply this to any image in no time!
Prompt… pic.twitter.com/TYhD7VzZAS
— Umesh (@umesh_ai) October 24, 2024
Here’s a bit more on how the new editing features work:
We’re testing two new features today: our image editor for uploaded images and image re-texturing for exploring materials, surfacing, and lighting. Everything works with all our advanced features, such as style references, character references, and personalized models pic.twitter.com/jl3a1ZDKNg
— Midjourney (@midjourney) October 23, 2024
I’ve become an Ideogram superfan, using it to create imagery daily, so I’m excited to kick the tires on this new interactive tool—especially around its ability to synthesize new text in the style of a visual reference.
Today, we’re introducing Ideogram Canvas, an infinite creative board for organizing, generating, editing, and combining images.
Bring your face or brand visuals to Ideogram Canvas and use industry-leading Magic Fill and Extend to blend them with creative, AI-generated content. pic.twitter.com/m2yjulvmE2
— Ideogram (@ideogram_ai) October 22, 2024
You can upload your own images or generate new ones within Canvas, then seamlessly edit, extend, or combine them using industry-leading Magic Fill (inpainting) and Extend (outpainting) tools. Use Magic Fill and Extend to bring your face or brand visuals to Ideogram Canvas and blend them with creative, AI-generated elements. Perfect for graphic design, Ideogram Canvas offers advanced text rendering and precise prompt adherence, allowing you to bring your vision to life through a flexible, iterative process.
Filmmaker & Pika Labs creative director Matan Cohen Grumi makes this town look way more dynamic than usual (than ever?) through the power of his team’s tech:
Took @pika_labs AI effects to the streets of San Jose. It’s crazy what you can create with just a phone, Pika and some basic edits #pikaffects pic.twitter.com/uzN2KyLHnh
— Matan Cohen-Grumi (@MatanCohenGrumi) October 19, 2024
Adobe’s new generative 3D/vector tech is a real head-turner. I’m impressed that the results look like clean, handmade paths, with colors that match the original—and not like automatic tracing of crummy text-to-3D output. I can’t wait to take it for a… oh man, don’t say it don’t say it… spin.
Oh man, for years we wanted to build this feature into Photoshop—years! We tried many times (e.g. I wanted this + scribble selection to be the marquee features in Photoshop Touch back in 2011), but the tech just wasn’t ready. But now, maybe, the magic is real—or at least tantalizingly close!
Being a huge nerd, I wonder about how the tech works, and whether it’s substantially the same as what Magnific has been offering (including via a Photoshop panel) for the last several months. Here’s how I used that on my pooch:

But even if it’s all the same, who cares?
Being useful to people right where they live & work, with zero friction, is tremendous. Generative Fill is a perfect example: similar (if lower quality) inpainting was available from DALL•E for a year+ before we shipped GenFill in Photoshop, but the latter has quietly become an indispensible, game-changing piece of the imaging puzzle for millions of people. I’d love to see compositing improvements go the same way.
As I drove the Micronaxx to preschool back in 2013, Macklemore’s “Can’t Hold Us” hit the radio & the boys flipped out, making their stuffed buddies Leo & Ollie go nuts dancing to the tune. I remember musing with Dave Werner (a fellow dad to young kids) about being able to animate said buddies.
Fast forward a decade+, and now Dave is using Adobe’s recently unveiled Firefly Video model to do what we could only dimly imagine back then:
Bringing stuffed animals to life with Adobe Firefly Generate Video. pic.twitter.com/XSbQxaIDiD
— Dave Werner (@okaysamurai) October 16, 2024
Time to unearth Leo & get him on stage at last. :->
Enjoy the latest from Magnific impresario Javi Lopez!
PART 3: Handed my vacation videos to an AI for auto editing, and now I’m pretty sure I’ll have nightmares for life pic.twitter.com/jYX1TZ4rMX
— Javi Lopez (@javilopen) October 12, 2024
As soon as Google dropped DreamBooth back in 2022, people have been trying—generally without much success—to train generative models that can incorporate the fine details of specific products. Thus far it just hasn’t been possible to meet most brands’ demanding requirements for fidelity.
Now tiny startup Flair AI promises to do just that—and to pair the object definitions with custom styling and even video. Check it out:
You can now generate brand-consistent video advertisements for your products on @flairAI_
1. Train a model on your brand’s aesthetic
2. Train a model on your clothing or product
3. Combine both models in one prompt
4. Animate✨In beta – comment/RT for access and free credits pic.twitter.com/88NYLVOFSQ
— Mickey Friedman (@mickeyxfriedman) October 7, 2024
I was super hyped last year when Meta announced “Emu Edit” tech for selectively editing images using just language:
Now you can try the tech via Meta.ai and in various apps:
Meta has casually released the best AI image editor
You can upload your image to Meta AI and just write the edits you want to make.
Accessible for free in WhatsApp, Instagram, Messenger, Facebook, etc. pic.twitter.com/jJEhMdJadT
— Paul Couvert (@itsPaulAi) October 2, 2024
In my limited experience so far, it’s cool but highly unpredictable. I’ll test it further, and I’d love to know how it works for you. Meanwhile you can try similar techniques via https://playground.com/:
Welcome to the new Playground
Use AI to design logos, t-shirts, social media posts, and more by just texting it like a person.
Watch: pic.twitter.com/eSwJcJUxtB
— Playground (@playground_ai) September 3, 2024
As always, I’m blown away in equal parts by:
Days of Miracles & Wonder, amirite?
Wow @runwayml just dropped an updated Gen-3 Alpha Turbo Video-to-Video mode & it’s awesome! It’s super fast & lets you do 9:16 portrait video. Anything is possible! pic.twitter.com/AxeFaJwAPR
— Blaine Brown (@blizaine) September 28, 2024
Oof. But of course he’s right that a tool is just a tool, not a provider of meaning & value unto itself.
GDT says it all here. pic.twitter.com/pK5WPtDY7l
— Todd Vaziri (@tvaziri) September 17, 2024
Chaos reigns!
I have no idea what AI and other tools were used here, but it’d be fun to get a peek behind the curtain. As a commenter notes,
The meandering strings in the soundtrack. The hard studio lighting of the close-ups. The midtone-heavy Technicolor grading. The macro-lens DOF for animation sequences. This is spot-on 50’s film aesthetic, bravo.
[Via Andy Russell]
And if that headline makes no sense, it probably just means your not terminally AI-pilled, and I’m caught flipping a grunt. 😉 Anyway, the tiny but mighty crew at Krea have brought the new Flux text-to-image model—including its ability to spell—to their realtime creation tool:
Flux now in Realtime.
available in Krea with hundreds of styles included.
free for everyone. pic.twitter.com/4gmMOmcUvg
— KREA AI (@krea_ai) September 12, 2024
What a fun little project & great NYC vibe-catcher: the folks at Runway captured street scenes with a disposable film camera, then used their model to put the images in motion. Check it out:
Shooting visual effects with a disposable camera and Gen-3 Alpha. pic.twitter.com/QRd3cI4Hqr
— Runway (@runwayml) September 6, 2024
I love seeing how scrappy creators combine tools in new ways, blazing trails that we may come to see as commonplace soon enough. Here Eric Solorio (enigmatic_e) shows how he used Viggle & other tools to create his viral Deadpool animation:
As promised, here is a breakdown of how I did the Deadpool animation I recently posted. pic.twitter.com/F130Skq17U
— enigmatic_e (@8bit_e) August 1, 2024
See also some of his luchador moves, plus more on his various feeds:
…with bears! Courtesy of image references in Photoshop GenFill:
— Anna McNaught (@annamcnaughty) August 28, 2024
I’ve been having a ball using the new Ideogram app for iOS to import photos & remix them into new creations. This is possible via their web UI as well, but there’s something extra magical about the immediacy of capture & remix. Check out a couple quick explorations I did while out with the kids, starting from a ballcap & the fuel tank of an old motorcycle:
More examples, riffing on a classic @TriumphAmerica fuel tank: pic.twitter.com/cZ5USqyGFN
— John Nack (@jnack) August 27, 2024
I love this level of transparency from the folks behind Photo AI. Developer @levelsio reports,
[Flux] made Photo AI finally good enough overnight to be actually used by people and be satisfied with the results… it’s more expensive [than SD] but worth it because the photos are way way better… Not sure about profitability but with SD it was about 85% profit. With Flux def less maybe 65%… Very unplanned and grateful the foundational models got better.
We’re arguably in something of a trough of disillusionment in the AI-art hype cycle, but this kind of progress gives reason for hope: more quality & more utility do translate into more sustainable value—and there’s every reason to think that things will only improve from here.
Flux, the new AI model, changes businesses (and lives)
It made https://t.co/1vEawpI5vb finally good enough overnight to be actually used by people and be satisfied with the results
All my improvements before helped but now it’s accelerating with Flux’s photo quality pic.twitter.com/BiAqi5BgnY
— @levelsio (@levelsio) August 21, 2024
Listen, I know that it’s a lot more seductive & cathartic to say “I f*cking hate generative AI,” and you can get 90,000+ likes for doing so, but—believe it or not—thoughtfulness & nuance actually matter. That is, how one uses generative tech can have very different implications for the creative community.
It’s therefore important to evaluate a range of risk/reward scenarios: What’s unambiguously useful & low-risk, vs. what’s an inducement to ripping people off, and what lies in the middle?
I see a continuum like this (click/tap to see larger):

None of this will draw any attention or generate much conversation—at least if my attempts to engage people on Twitter are any indication—but it’s the kind of thing actual toolmakers must engage with if we’re to make progress together. And so, back to work.
PS—This, always this:
This kind of foolishness soothes my soul. :-p
Some Chinese dudes imitating AI videos lol this is next level pic.twitter.com/LqB3O327Kr
— GioM (@theGioM) August 15, 2024
My friend Nathan has fed a mix of Schwarzenegger photos & drawings from Aesop’s Fables into the new open-source Flux model, creating a rad woodcut style. That’s interesting enough on its own—but it’s so 24 hours ago, and thus he’s now taken to animating the results. Check out the thread below for details:
Animating yesterday’s #FLUX woodcut Arnold using one of my favorite clips from the old soundboards
This uses Follow-Your-Emoji / Reference UNet in ComfyUI, which did a better job than LivePortrait.
Some comparison results in thread #aivideo pic.twitter.com/C9pgWgVJS5
— Nathan Shipley (@CitizenPlain) August 15, 2024