My family, having seen so many of my AI-powered image generations over the last 3 years, is just utterly inured to them. So, for my MiniMe’s 16th, I sketched up the patriotic little HO-scale engine we’re getting him, along with a cute large ground squirrel (to quote the Dude, “Nice marmot”).
I feel like this is my micro version of when the world revolted against too-perfect Instagram culture, swinging towards Snapchat & stories, where “rough is real,” and flaws are a feature. In any case, my dude was happy as a clam—and that’s all that matters to me.
For my son’s birthday, I ditched AI and broke out my pen. Felt good to work without a net. pic.twitter.com/TJedYCPWMj
Okay, so this isn’t precisely what I thought it was at first (video inpainting), but rather an creation->inpainting->animation flow. Still, the results look impressive:
How it works:
→ Generate an image in Higgsfield Soul → Inpaint directly with a mask and a prompt → Combine with Camera moves, VFX, and Avatars to turn static edits into living, speaking visuals pic.twitter.com/ENHqdA3WHm
“If you’re into weird cars, forgotten history, and stories that don’t end well, hit that subscribe button.”
I found this piece really interesting, not least because my wife & I are headed to Africa for the first time next week, and I’m eager to learn what kinds of vehicles & roads we’ll experience. Seems like something like the Africar would make a ton of sense in many places:
As I’ve noted previously, Google has been trying to crack the try-on game for a long time. Back in the day (c. 2017), we really want to create AR-enabled mirrors that could do this kind of thing. The tech wasn’t quite ready, and for the realtime mirror use case it likely still isn’t, but check out the new free iOS & Android app Doppl:
In May, Google Shopping announced the ability to virtually try billions of clothing items on yourself, just by uploading a photo. Doppl builds on these capabilities, bringing additional experimental features, including the ability to use photos or screenshots to “try on” outfits whenever inspiration strikes.
Doppl also brings your looks to life with AI-generated videos — converting static images into dynamic visuals that give you an even better sense for how an outfit might feel. Just upload a picture of an outfit, and Doppl does the rest.
Several years ago, MyHeritage saw a huge (albeit short-lived) spike in interest from their Deep Nostalgia feature that animated one’s old photos. Everything old is new again, in many senses. Check out Reddit founder Alexis Ohanian talk about how touching he found the tech—as well as tons of blowback from people who find it dystopian.
Damn, I wasn’t ready for how this would feel. We didn’t have a camcorder, so there’s no video of me with my mom. I dropped one of my favorite photos of us in midjourney as ‘starting frame for an AI video’ and wow… This is how she hugged me. I’ve rewatched it 50 times. pic.twitter.com/n2jNwdCkxF
I’ve heard people referring to the recent release of Google’s Veo 3 as the ChatGPT moment for video generation—that is, a true inflection point at which a mere curosity becomes something of real value. The spatial & character coherence of its output, and especially its ability to generate speech & other audio, turn it into a genuine storytelling tool.
You’ve probably seen some of the myriad vlogger-genre creations making the rounds. Here’s one of my faves:
I’ll note the fact of AI having been involved only because at this point who cares whether AI was involved? We’re happily reaching a plane of maturity where the particular mix of tooling is much less interesting than the vision & vibe.
John Gruber recently linked back to this clip in which designer Neven Mrgan highlights what feels like an important consideration in the age of mass-generated AI “designs”:
I think that was what mattered is that they looked rich, they looked like a lot of work had been put into them. That’s what people latch onto. It seems it’s something that, yes, they should have spent money on, and they should be spending time on right now.
Regardless of what tools were used in the making of a piece, does it feel rich, crafted, thoughtfully made? Does it have a point, and a point of view? As production gets faster, those qualities will become all the more critical for anything—and anyone—wishing to stand out.
This could be an awesome opportunity for the right person, who’d get to work on things I’ve wanted the team to do for 15+ years!
We’re looking for an expert technical product manager to lead Photoshop’s foundational architecture and performance strategy. This is a pivotal role responsible for evolving the core technologies that power Photoshop’s speed, stability, and future scalability across platforms.
You’ll drive major efforts to modernize our rendering and compute architecture, migrate legacy systems to more scalable platforms, and accelerate performance through GPU and hardware optimization. This work touches nearly every part of Photoshop, from canvas rendering to feature responsiveness to long-term cross-platform consistency.
This is a principal-level individual contributor role with the potential to grow a team in the future.
I interviewed many hundreds of PM candidates at Google, and if things were going well, I’d ask, “Tell me about a product you hate that you use regularly. Why do you hate it?”
This proved to be a great bozo detector. Does this person have curiosity, conviction, passion, unreasonableness? Were they forced into coding & now just want to escape life in the damn debugger, or do they have a semi-pathological need to build stuff they’re proud of? Would I want them in the proverbial foxhole with me? Are they willing to sweep the floor?
Unsurprisingly, most candidates offer shallow, banal answers (“Uh, wow… I mean, I guess the ESPN app is kinda slow…?”), whereas great ones explain not just what sucks, but why it sucks. Like, why—systemically—is every car infotainment system such crap? Those are the PMs I want asking the questions, then questioning the answers.
——-
Specifically the car front, as Tolstoy might say, “Each one is unhappy in its own way.” The most interesting thing, I think, isn’t just to talk about the crappy mismatched & competing experiences, but rather about why every system I’ve ever used sucks. The answer can’t be “Every person at every company is a moron”—so what is it?
So much comes down to the structure of the industry, with hardware & software being made by a mishmash of corporate frenemies, all contending with a soup of regulations, risk aversion (one recall can destroy the profitability of a whole product line), and surprisingly bargain-bin electronics.
Check out this short vid for some great insights from Ford CEO Jim Farley:
Ford CEO Jim Farley on why it’s so difficult for legacy car companies to get software right & why @Tesla’s vertically integrated approach is the right one:
“We farmed out all the modules that control the vehicles to our suppliers because we could bid them against each other, so… pic.twitter.com/kWsIaiOJlI
A while back, Sam Harris & Ricky Gervais discussed the impossibility of translating a joke discovered during a dream (“What noise does a monster make?”) back into our consensus waking reality. Like… what?
I get the same vibes watching ChatGPT try to dredge up some model of me and of… humor?… in creating a comic strip based on our interactions. I find it uncanny, inscrutable, and yet consequently charming all at once.
“Hey ChatGPT, based on what you know about me, please create a four-panel comic you think I’d like…” https://t.co/U7WRfShGRh
Splice (2D/3D design in your browser) has added support for progressive blur & gradients, and the results look awesome.
I haven’t seen anything advance like this in Adobe‘s core apps in maybe 20 years— maybe 25, since Illustrator & Acrobat added support for transparency.
We are adding Progressive Blur + Gradients to Hana! All interactive, all real-time.
On an aesthetically similar note, check out the launch video for the new version of Sketch (still very much alive & kicking in an age of Figma, it seems):
Remember when we said auto layout was coming to Sketch? It’s here. It’s called Stacks, and it’s part of our biggest release ever — out now.
There’s a lot to cover, so buckle up and we’ll give you a tour.
Also, stick around for a surprise at the end of the thread
Opt in to get started: Head over to Search Labs and opt into the “try on” experiment.
Browse your style: When you’re shopping for shirts, pants or dresses on Google, simply tap the “try it on” icon on product listings.
Strike a pose: Upload a full-length photo of yourself. For best results, ensure it’s a full-body shot with good lighting and fitted clothing. Within moments, you can see how the garment will look on you.
Several years ago, my old teammates shared some promising research on how to facilitate more interesting typesetting. Check out this 1-minute overview:
Ever since the work landed in Adobe Express a while back, I’ve wondered why it hadn’t yet made its way to Photoshop or Illustrator. Now, at least, it looks like it’s on its way to PS:
The feature looks cool, and I’m eager to try it out, but I hope that Adobe will keep trying to offer something more semantically grounded (i.e. where word size is tied to actual semantic importance, not just rectangular shape bounds)—like what we shipped last year:
Man, for 18 years (yes, I keep the receipts) I’ve been wanting to ship an interactive relighting experience—and now my team has done it! Check out the quick demo below plus details on DP Review.
Good news! You too can capture footage exactly like this. You just need a $100,000 Phantom Flex 4K with a Canon 50-1000mm lens—oh, and you need to be hanging out the side of a Black Hawk helicopter:
We’ve released the code for LegoGPT. This autoregressive model generates physically stable and buildable designs from text prompts, by integrating physics laws and assembly constraints into LLM training and inference.
I have to admit, I don’t know Erwitt’s photography nearly as well as I know his name, but this largely humorous new collection makes me want to change that:
Continuing their excellent work to offer more artistic control over image creation, the fast-moving crew at Krea has introduced GPT Paint—essentially a simple canvas for composing image references to guide the generative process. You can directly sketch, and/or position reference images, then combine the input with prompts & style references to fine-tune compositions:
introducing GPT Paint.
now you can prompt ChatGPT visually through edit marks, basic shapes, notes, and reference images.
Historically, approaches like this have sounded great but—at least in my experience—have fallen short.
Think about what you’d get from just saying “draw a photorealistic beautiful red Ferrari” vs. feeing in a crude sketch + the same prompt.
In my quick tests here, however, providing a simple reference sketch seems helpful—maybe because GPT-4o is smart enough to say, “Okay, make a duck with this rough pose/position—but don’t worry about exactly matching the finger-painted brushstrokes.” The increased sense of intentionality & creative ownership feels very cool. Here’s a quick test:
I’m not quite sure where the spooky skull and, um, lightning-infused martini came from. 🙂
Director John Likens and FX Supervisor Tomas Slancik dissect existential collapse in Your Friends & Neighbors’ haunting opener, blending Jon Hamm’s live-action gravitas with a symphony of digital decay. […]
Shot across two days and polished by world-class VFX artists, the title sequence mirrors Hamm’s crumbling protagonist, juxtaposing his stoic performance against hyper-detailed destruction.
Having created 200+ images in just the last month via this still-new image model (see new blog category that gathers some of them), I’m delighted to say that my team is working to bring it to Microsoft Designer, Copilot, and beyond. From the boss himself:
5/ Create: This one is fun. Turn a PowerPoint into an explainer video, or generate an image from a prompt in Copilot with just a few clicks.
We’ve also added new features to make Copilot even more personalized to you, plus a redesigned app built for human-agent collaboration. pic.twitter.com/m1oTf53aai
“You’re now the proud owner of the most dangerously cozy footwear in the sky. Plush, cartoon A-10 Warthogs with big doe eyes and turbine engines ready to warm your toes and deliver cuddly close air support. Let me know if you want tiny GAU-8 Gatling gun detailing on the front.”… pic.twitter.com/lKLRJGALaw
Back at Adobe we introduced Firefly text-to-vector creation, but behind the scenes it was really text-to-image-to-tracing. That could be fine, actually, provided that the conversion process did some smart things around segmenting the image, moving objects onto their own layers, filling holes, and then harmoniously vectorizing the results. I’m not sure whether Adobe actually got around to shipping that support.
In any event, StarVector promises actual, direct creation of SVG. The results look simple enough that it hasn’t yet piqued my interest enough to spend my time with it, but I’m glad that folks are trying.
I really hope that the makers of traditional vector-editing apps are paying attention to rich, modern, GPU-friendly techniques like this one. (If not—and I somewhat cynically expect that it’s not—it won’t be for my lack of trying to put it onto their radar. ¯\_(ツ)_/¯)
Introducing Vector Feathering — a new way to create vector glow and shadow effects. Vector Feathering is a technique we invented at Rive that can soften the edge of vector paths without the typical performance impact of traditional blur effects. (Audio on) pic.twitter.com/39kfjmFsTJ
I know only what you see below, but Magic Animator (how was that domain name available?) promises to “Animate your designs in seconds with AI,” which sounds right up my alley, and I’ve signed up for their waitlist.
Three years ago (seems like an eternity), I remarked regarding generative imaging.,
The disruption always makes me think of The Onion’s classic “Dolphins Evolve Opposable Thumbs“: “Holy f*ck, that’s it for us monkeys.” My new friend August replied with the armed dolphin below.
I’m reminded of this seeing Google’s latest AI-powered translation (?!) work. Just don’t tell them about abacuses!
Meet DolphinGemma, an AI helping us dive deeper into the world of dolphin communication. pic.twitter.com/2wYiSSXMnn
Wait, first, WTF is MCP? Check out my old friend (and former Illustrator PM) Mordy’s quick & approachable breakdown of Model Context Protocol and why it promises to be interesting to us (e.g. connecting Claude to the images on one’s hard drive).