Check out this recording of evangelist Paul Trani’s 1-hour deep dive into Firefly, including examples of how to refine & extend its output in Photoshop:
I enjoyed hearing my colleagues & outside folks discussing the origin, vision, and road ahead for Adobe Firefly in this livestream…
Eric Snowden is the VP of Design at Adobe and is responsible for the product design teams for the Digital Media business, which include Creative Cloud…. Nishat Akhtar is a designer and creative leader with 15+ years of experience in designing and leading initiatives for global brands… Danielle Morimoto is a Design Manager for Adobe Express, based in San Francisco.
…and this Twitter space, featuring our group’s CTO Ely Greenfield, along with creator Karen X. Cheng (whose work I’ve featured here countless times), illustrator & brush creator Kyle T. Webster, and director of design Samantha Warren. Scrub ahead to about 2:45 to get to the conversation.
Made with genuine diabeetus! All right stop, collaborate and listen:
On one hand, you may be convinced we somehow assembled the original cast of The Matrix alongside the ghost of Wilford Brimley to record one of the greatest rap covers of all time. On the other hand, you may find it more believable that we’ve been experimenting with AI voice trainers and lip flap technology in a way that will eventually open up some new doors for how we make videos. You have to admit, either option kind of rules.
Hey, remember when we launched Adobe Firefly what feels like 63 years ago? 😅 OMG, what a week. I am so tired & busy trying to get folks access (thanks for your patience!), answer questions, and more that I’ve barely had time to catch up on all the great content folks are making. I’ll work on that soon, and in the meantime, here are three quick clips that caught my eye.
First, OG author Deke McClelland shows off type effects:
I really appreciate hearing Karen X. Cheng’s thoughts on the essential topics of consent, compensation, and more. We’ve been engaging in lots of very helpful conversations with creators, and there’s of course much more to sort through. As always, your perspective here is most welcome.
I’m so pleased—and so tired! 😅—to be introducing Adobe Firefly, the new generative imaging foundation that a passionate band of us have been working to bring to the world. Check out the high-level vision…
…as well as the part more directly in my wheelhouse: the interactive preview site & this overview of great stuff that’s waiting in the wings:
I’ll have a lot more to share soon. In the meantime, we’d love to hear what you think of what you see so far!
Starting today our community can test Midjourney V5. It has much higher image quality, more diverse outputs, wider stylistic range, support for seamless textures, wider aspect ratios, better image prompting, wider dynamic range and more. Let’s explore!
I’ve spent the last ~year talking about my brain being “DALL•E-pilled,” where I’ve started seeing just about everything (e.g. a weird truck) as some kind of AI manifestation. But that’s nothing compared to using generative imaging models to literally see your thoughts:
Researchers Yu Takagi and Shinji Nishimoto, from the Graduate School of Frontier Biosciences at Osaka University, recently wrote a paper outlining how it’s possible to reconstruct high res images (PDF) using latent diffusion models, by reading human brain activity gained from functional Magnetic Resonance Imaging (fMRI), “without the need for training or fine-tuning of complex deep generative models” (via Vice).
Check out this integration of sketch-to-image tech—and if you have ideas/requests on how you’d like to see capabilities like these get more deeply integrated into Adobe tools, lay ’em on me!
Also, it’s not in Photoshop, but as it made me think of the Photo Restoration Neural Filter in PS, check out this use of ControlNet to revive an old family photo:
I’m really excited to see what kinds of images, not to mention videos & textured 3D assets, people will now be able to generate via emerging techniques (depth2img, ControlNet, etc.):
In a demo video, Qualcomm shows version 1.5 of Stable Diffusion generating a 512 x 512 pixel image in under 15 seconds. Although Qualcomm doesn’t say what the phone is, it does say it’s powered by its flagship Snapdragon 8 Gen 2 chipset (which launched last November and has an AI-centric Hexagon processor). The company’s engineers also did all sorts of custom optimizations on the software side to get Stable Diffusion running optimally.
This new capability in Stable Diffusion (think image-to-image, but far more powerful) produces some real magic. Check out what I got with some simple line art:
My friend Bilawal Sidhu made a 3D scan of his parents’ home (y’know, as one does), and he recently used the new ControlNet functionality in Stable Diffusion to restyle it on the fly. Check out details in this post & in the vid below:
1992 Pink Floyd laser light show in Dubuque, IA—you are back. 😅
Through this AI DJ project, we have been exploring the future of DJ performance with AI. At first, we tried to make an AI-based music selection system as an AI DJ. In the second iteration, we utilized a few AI models on stage to generate real-time symbolic music (i.e., MIDI). In the performance, a human DJ (Tokui) controlled various parameters of the generative AI models and drum machines. This time, we aim to advance one step further and deploy AI models to generate audio on stage in near real-time. Everything you hear during the performance will be pure AI-generation (no synthesizer, no drum machine).
In this performance, Emergent Rhythm, the human DJ will become an AJ or “AI Jockey” instead of a Disk Jockey, and he is expected to tame and ride the AI-generated audio stream in real-time. The distinctive characteristics of AI-based audio generation and “morphing” will provide a unique and even otherworldly sonic experience for the audience.
Introducing the new DigitalFUTURES course of free AI tutorials.
Several of the top AI designers in the world are coming together to offer the world’s first free, comprehensive course in AI for designers. This course starts off at an introductory level and gets progressively more advanced. 18 Feb, Introductory Session 10.00 am EST, 4.00 pm CET, 11.00 pm China What is AI? What are Midjourney, DALL•E, Stable Diffusion, etc.? What is GPT3? What is ChatGPT? And how are they revolutionizing design?
Check out this craziness (which you can try online) from Google researchers, who write, “We introduce Piano Genie, an intelligent controller that maps 8-button input to a full 88-key piano in real time”:
Paul Trillo used Runway’s new Gen-1 experimental model to create a Cubist Simpsons intro:
“The Simpsons” but make it an experimental cubist stop motion. Felt right given long tradition of reanimating the Simpsons intro. Created with the spellbindingly addictive #Gen1 AI video generator from @runwayml — still early days but step in the future if #animation#ai#aiartpic.twitter.com/XZLgGpBLCw
Check out this new generative stylization model. I’m intrigued by the idea of using simple primitives (think dollhouse furniture) to guide synthesis & stylization (e.g. of the buildings shown briefly here).
Today, Generative AI takes its next big step forward.
Introducing Gen-1: a new AI model that uses language and images to generate new videos out of existing ones.
Last month Paul Trillo shared some wild visualizations he made by walking around Michelangelo’s David, then synthesizing 3D NeRF data. Now he’s upped the ante with captures from the Louvre:
NeRFs of the Louvre made from a handful of short iPhone videos shot during a location scout last month. Each shot reimagined over a month later. The impossibilities are endless. More to come…
I got my professional start at AGENCY.COM, a big dotcom-era startup co-founded by creative whirlwind Kyle Shannon. Kyle has been exploring AI imaging like mad, and recently he’s organized an AI Artists Salon that anyone is welcome to join in person (Denver) or online:
The AI Artists Salon is a collaborative group of creatively-minded people and we welcome anyone curious about the tsunami of inspiring generative technologies already rocking our our world. See Community Links & Resources.
On Tuesday evening I had the chance to present some ideas & progress that has inspired me—nothing confidential about Adobe work, of course, but hopefully illuminating nonetheless. If you’re interested, check it out (and pro tip: if you set playback to 1.5x speed or higher, I sound a lot sharper & funnier!).
Here’s an example made from a quick capture I did of my friend (nothing special, but amazing what one can get simply by walking in a circle while recording video):
As luck (?) would have it, the commercial dropped on the third anniversary of my former teammate Jon Barron & collaborators bringing NeRFs into existence:
Three years ago today, the project that eventually became NeRF started working (positional encoding was the missing piece that got us from "hmm" to "wow"). Here's a snippet of that email thread between Matt Tancik, @_pratul_, @BenMildenhall, and me. Happy birthday NeRF! pic.twitter.com/UtuQpWsOt4
Thank God for the vibrant developer community—esp. Adobe vet Christian Cantrell (who somehow finds time to rev his plugin while serving as VP of product for Stability.ai):
“HEY MAN, you ever drop acid?? No? Well I do, and it looks *just like this*!!” — an excitable Googler when someone wallpapered a big meeting room in giant DeepDream renderings
In a similar vein, have fun tripping balls with AI, courtesy of Remi Molettee:
The company has announced a new mode for their Canvas painting app that turns simple brushstrokes into 360 environment maps for use in 3D apps or Omniverse. Check out this quick preview:
My teammate CJ Gammon has released a handy new Chrome extension that lets you select any image, then use it as the seed for new image generation. Check it out:
In this beautiful work from Paul Trillo & co., AI extends—instead of replaces—human creativity & effort:
Here’s a peek behind the scenes:
This project would have never existed without the use of AI. A variety of tools were used from #dalle2 and #stablediffusion to generate the background assets Automatic1111 #img2img and @runwayml to process the video along with @AdobeAE to create the camera moves and transitions pic.twitter.com/FwqwWto966
1. Take reference photo (you can use any photo – e.g. your real house, it doesn’t have to be dollhouse furniture) 2. Set up Stable Diffusion Depth-to-Image (google “Install Stable Diffusion Depth to Image YouTube”) 3. Upload your photo and then type in your prompts to remix the image
We recommend starting with simple prompts, and then progressively adding extra adjectives to get the desired look and feel. Using this method, @justinlv generated hundreds of options, and then we went through and cherrypicked our favorites for this video
I’m not sure what to say about “The first rap fully written and sung by an AI with the voice of Snoop Dogg,” except that now I really want the ability to drop in collaborations by other well known voices—e.g. Christopher Walken.
Maybe someone can now lip-sync it with the faces of YoDogg & friends:
The marketers at Heinz had a little fun noticing that an AI image-making app (DALL•E, I’m guessing) tended to interpret requests for “ketchup” in the style of Heinz’s iconic bottle. Check it out:
The whole community of creators, including toolmakers, continues to feel its way forward in the fast-moving world of AI-enabled image generation. For reference, here are some of the statements I’ve been seeing:
“Kickstarter must, and will always be, on the side of creative work and the humans behind that work. We’re here to help creative work thrive.”
Key questions they’ll ask include “Is a project copying or mimicking an artist’s work?” and “Does a project exploit a particular community or put anyone at risk of harm?”
From 3dtotal Publishing:
“3dtotal has four fundamental goals. One of them is to support and help the artistic community, so we cannot support AI art tools as we feel they hurt this community.”
“We oppose the commercial use of Artificially manufactured images and will not allow AI into our annual competitions at all levels.”
“AI was trained using copyrighted images. We will oppose any attempts to weaken copyright protections, as that is the cornerstone of the illustration community.”
This stuff—creating 3D neural models from simple video captures—continues to blow my mind. First up is Paul Trillo visiting the David:
Finally got to see Michaelangelo's David in Florence and rather than just take a photo like normal person, I spent 20 minutes walking around it capturing every angle looking like an insane person. It's hard to look cool when making a #NeRF but damn it looks cool later @LumaLabsAIpic.twitter.com/sLGJ2CKCJy
Numerous apps are promising pure text-to-geometry synthesis, as Luma AI shows here:
✨ Introducing Imagine 3D: a new way to create 3D with text! Our mission is to build the next generation of 3D and Imagine will be a big part of it. Today Imagine is in early access and as we improve we will bring it to everyone https://t.co/VIdilw7kpapic.twitter.com/v6Yi0mwZsY
On a more immediately applicable front, though, artists are finding ways to create 3D (or at least “two-and-a-half-D”) imagery right from the output of apps like Midjourney. Here’s a quick demo using Blender:
In a semi-related vein, I used CapCut to animate a tongue-in-cheek self portrait from my friend Bilawal:
Creative Reality Studio from D-ID (the folks behind the MyHeritage Deep Nostalgia tech that blew up a couple of years ago) can generate faces & scripts, then animate them. I find the results… interesting?
I believe strongly that creative tools must honor the wishes & rights of creative people. Hopefully that sounds thuddingly obvious, but it’s been less obvious how to get to a better state than the one we now inhabit, where a lot of folks are (quite reasonably, IMHO) up in arms about AI models having been trained on their work, without their consent. People broadly agree that we need solutions, but getting to them—especially via big companies—hasn’t been quick.
Thus it’s great to see folks like Mat Dryhurst & Holly Herndon driving things forward, working with Stability.ai and others to define opt-out/-in tools & get buy-in from model trainers. Check out the news:
Artist & musician Ben Morin has been making some impressive pop-culture mashups, turning well-known characters into babies (using, I believe, Midjourney to combine a reference image with a prompt). Check out the results.
Our friend Christian Cantrell (20-year Adobe vet, now VP of Product at Stability.ai) continues his invaluable world to plug the world of generative imaging directly into Photoshop. Check out the latest, available for free here:
1) Support for extreme resolutions (up to 1MP). 2) Automatic selection of optimal models. 3) Access to all SD versions (1.4, 1.5, 2.0, and 2.1). 4) Account credits and avatar.https://t.co/gqFWpAkfnopic.twitter.com/DSwbC2xstL
It’s insane to me how much these emerging tools democratize storytelling idioms—and then take them far beyond previous limits. Recently Karen X. Cheng & co. created some wild “drone” footage simply by capturing handheld footage with a smartphone:
Now they’re creating an amazing dolly zoom effect, again using just a phone. (Click through to the thread if you’d like details on how the footage was (very simply) captured.)
NeRF update: Dollyzoom is now possible using @LumaLabsAI I shot this on my phone. NeRF is gonna empower so many people to get cinematic level shots Tutorial below –
Check out the latest magic, as described by Gizmodo:
To make an age-altering AI tool that was ready for the demands of Hollywood and flexible enough to work on moving footage or shots where an actor isn’t always looking directly at the camera, Disney’s researchers, as detailed in a recently published paper, first created a database of thousands of randomly generated synthetic faces. Existing machine learning aging tools were then used to age and de-age these thousands of non-existent test subjects, and those results were then used to train a new neural network called FRAN (face re-aging network).
When FRAN is fed an input headshot, instead of generating an altered headshot, it predicts what parts of the face would be altered by age, such as the addition or removal of wrinkles, and those results are then layered over the original face as an extra channel of added visual information. This approach accurately preserves the performer’s appearance and identity, even when their head is moving, when their face is looking around, or when the lighting conditions in a shot change over time. It also allows the AI generated changes to be adjusted and tweaked by an artist, which is an important part of VFX work: making the alterations perfectly blend back into a shot so the changes are invisible to an audience.
As I say, another day, another specialized application of algorithmic fine-tuning. Per Vice:
For $19, a service called PhotoAI will use 12-20 of your mediocre, poorly-lit selfies to generate a batch of fake photos specially tailored to the style or platform of your choosing. The results speak to an AI trend that seems to regularly jump the shark: A “LinkedIn” package will generate photos of you wearing a suit or business attire…
…while the “Tinder” setting promises to make you “the best you’ve ever looked”—which apparently means making you into an algorithmically beefed-up dudebro with sunglasses.
Meanwhile, the quality of generated faces continues to improve at a blistering pace:
…thus inducing fans to reply with their own variations (click tweet above to see the thread). Among the many fun Snoop Doggs (or is it Snoops Dogg?), I’m partial to Cyberpunk…