Welcome to the Friday edition of our newsletter. We spend Fridays going deeper into tools and trends related to generative AI (and Tuesdays sharing news updates). This week Professor Porter discusses tips for using computer vision in a conversation with an LLM.

Computer vision in a conversation?

Over the past few Fridays, we’ve tried to provide practical tips for enhancing your use of generative AI tools, focusing on AI flyers that avoid the generic ChatGPT style that’s been going around, and when to disclose your usage of AI. In keeping with that practical trend, this week we’ll look at how to make the most of computer vision in conversation with large language models.

To get started, if you have an image on your computer or phone that you’d like to discuss with an LLM of your choice, you first need to load it into the conversation. This is easy to do: just click on the plus sign on the side of the text bar and look for the paperclip icon next to a phrase like “Upload files” or “Add photos & files.” From there, you’ll be prompted to select the desired image from your computer.

Alternatively, if you’re using an LLM tool on your smartphone, you can turn on voice mode and your camera and have a conversation about what your phone is seeing in real time. Or, you could simply use Google Lens on your smartphone.

Now that our desired visual information has entered the chat, what can we do with it?Read on below for eight ways you might proceed.

Free AI Update Aug. 26

Sign up for our free Back to School AI Update on Aug. 26. We will walk you through the latest news in the world of AI - from the government reviewing models to Work/Cowork to the Pope’s warning on AI and more. This event is offered free thanks to a sponsorship from technology and management consultancy, Lean TECHniques.

Want more AI? Sign up for our online classes

Want your brand in front of AI-forward professionals?

We're looking for a sponsor for our Fall AI Lunch Club – four virtual events that draw workers from Central Iowa and across the country who are actively figuring out how to use AI in their jobs. That's a specific, engaged audience you can't reach with a generic ad buy. Get the details here – or just reply to this email with questions.

1. Answering “What is it?” questions

We’ve all been in a situation where we’re confronted with an unknown object with no way to decipher what it is…until now. Simply ask the LLM to identify the object and it will likely get the answer right. Examples include:

  • Is this poison ivy?

  • What kind of bug is this?

  • What is this adapter for?

  • Is this fruit ripe enough to eat?

  • Is this a weed or a flower?

  • What kind of pill is this?

As an example from my own experience, I once had ChatGPT in voice mode identify a piece of rusted farm equipment that I encountered in a pasture in rural Oklahoma (I thought it might be some sort of furrowing device, but it proved to be a horse-drawn grass cutter).

Of course, identification should not be treated as infallible, especially when health or safety is involved. You probably should not eat a mushroom, handle a snake, or take a pill solely because an LLM told you what it was. But as a starting point for further investigation, visual identification can be remarkably useful.

2. Inquiring about visual information

More generally, you can do much more than ask “What is this?” of an image you upload to an LLM or show to it in real time. For instance, you might ask things like:

  • Where is the engine air filter located?  

  • How do I turn this thing off?

  • What are some possible sources of this leak?

  • What dish can I make with these ingredients?

  • What is the style of this painting?

Just last week I witnessed an Ace Hardware worker ask the Grok mobile app whether the drill bit in a countersink I wanted to purchase was removable. It wasn’t, so I didn’t buy it.

This is one of the most practical features of computer vision: You can point the camera at what you are seeing and ask questions in ordinary language rather than trying to describe the object, find the correct terminology, and search for the answer yourself.

3. Editing and restoring existing images

We know that tools like ChatGPT and Gemini can create images, but you can also load existing images as the subject of your conversation. In particular, you can edit your images, as I did when I removed an unwanted hand in a family photo that was going to be added to our yearly Christmas card. Here’s the before and after. Can you spot the difference?

Before…

…and after

A similar trick is to load an old photo and ask the LLM to restore it. You can remove scratches, sharpen blurry details, repair damaged sections, adjust lighting, and even colorize black-and-white photographs.

Earlier this year Snider shared in a newsletter an example of this very technique with a photo of his dad:

As with any AI-generated image, the restored version may introduce details that were not actually present in the original. It is better understood as an informed reconstruction than a perfect recovery of lost visual information.

4. Visualizing the transformation of physical spaces

Similar to the previous example, you can load a picture of a physical space and ask the LLM to transform the image in some way. You might ask it to repaint a room, replace flooring, redesign a landscape, add furniture, remove clutter, or show what a renovation might look like.

In our generative AI class, Snider and I ask our students to find an empty room on a real estate listing and have an LLM stage the room with furniture. One of my favorite examples from this past semester comes from a student who asked an LLM to decorate the room below on the left using the furnishings in the room below on the right (a painting by artist Richard Nadler as part of his HabiTextures series, which I discovered using Google Lens).

Here’s the finished product:

On the home front, ChatGPT helped my wife visualize a stone path along the side of our house, which she then built (based on a fairly accurate list of materials for the project that ChatGPT provided). Take a look at the progression below:

Uploaded image on the left, ChatGPT’s rendering in the middle, and the finished path on the right.

AI-generated renderings will not always preserve exact dimensions, structural details, or architectural constraints. But they can make an abstract idea concrete enough to discuss, revise, price, and potentially build.

5. Digitizing handwritten text

This example is straightforward: Just load an image of handwritten text and ask the LLM to digitize it for you. Immediate examples include converting handwritten class notes or notes from a meeting into text that can be pasted into a Word document. In our consulting work, Snider and I often fill up a whiteboard with ideas from a brainstorming session with a client, which can easily be converted into digital text for later processing.

6. Converting hand-drawn markups to coded objects

Along similar lines, we can draw up a diagram, the layout of a website, a flowchart, a slide template, or a wireframe for an app, load it to an LLM, and ask it to code it for us.

The resulting code might produce a functioning webpage, an interactive prototype, a digital flowchart, or a slide that follows the structure of the original sketch. Instead of starting with formal design software, you can begin with a pencil and a piece of paper.

This lowers the barrier between having an idea and building a working version of it. The initial output will usually need refinement, but the hand-drawn sketch can serve as a surprisingly effective first draft of the instructions.

7. Interpreting complex data visualizations and diagrams

LLMs with computer vision can also help users make sense of charts, technical diagrams, maps, dashboards, and other dense visual materials.

You might ask an LLM to:

  • Identify the components in a circuit-board diagram

  • Explain what the axes and lines in a graph represent

  • Summarize the main takeaway from a business dashboard

  • Trace the steps in a process diagram

  • Explain the symbols used on a map

  • Identify a suspicious pattern in a table or chart

  • Compare two different data visualizations

  • Translate a technical diagram into plain language

This can be especially useful when the visualization comes from a field outside your expertise. An LLM might not replace an engineer, statistician, or subject-matter expert, but it can help you understand what you are looking at and formulate better follow-up questions.

8. Providing text labels of images for visually impaired users

Computer vision can help make visual content more accessible by generating alt text and longer image descriptions for people who are blind or have low vision.

Simply upload an image and ask the LLM to describe it. You can specify where the description will appear and how much detail is appropriate. For example, you might request:

  • Concise alt text for a website

  • A detailed description of a chart or infographic

  • An image description for a social media post

  • A description of the important visual elements on a presentation slide

  • Alt text that focuses on information relevant to the surrounding article

This is especially useful when a document, website, presentation, or social media post contains many images that would otherwise need to be described individually. Of course, always be sure to double-check the LLM’s work.

Seeing LLMs differently

Most people still think of LLMs primarily as tools for generating and processing text. But once a model can see, the camera becomes another form of input, and the range of possible use cases expands considerably.

Instead of describing a problem, you can show it. Instead of translating a visual idea into technical instructions, you can sketch it. Instead of searching for the right terminology, you can point the camera and ask a question.

The results are not always perfect, and higher-stakes use cases require appropriate caution. But computer vision transforms an LLM from something that merely responds to what we type into something that can help us interpret, manipulate, and act upon the visual world around us.

Keep Reading