Disclosure: some links in this post are affiliate links. If you sign up through them I may earn a commission, at no extra cost to you.
Here’s something new: GoHighLevel bots can actually see the photos your customers send. Not just receive them — analyze the damage, read the context, and respond with real detail. This GoHighLevel tutorial covers the full setup: you build a prompt-based bot with personality, goal, and info sections that explain the image capability, wire it to a chat widget automation that turns the bot on when someone replies, and test it with real photos until the answers hold up.
Why does this matter for GoHighLevel bots?
It matters because bots could previously receive a photo but not understand it, leaving customers ignored or stuck repeating themselves. Now the bot reads the image, responds with real context, and can act on what it sees — analyzing damage, guiding the conversation, and moving toward a booked appointment without manual follow-up.
Here’s the problem this actually solves. Before this feature, a customer could send a photo into your chat and your bot had no idea what to do with it. No response, or worse, a repeated question that made it obvious nobody was paying attention. The customer thinks they’re talking to a human, gets ignored, and you end up doing the follow-up manually anyway.
Now the bot can read the image, respond with context, and act on what it sees. It’s not a gimmick — it’s your bot finally doing the one thing text alone could never do: look at a photo and react to it like it actually understands what’s in it.
What channels and file types work right now
As of this recording, image support works across live chat, Facebook Messenger, Instagram DMs, WhatsApp, SMS, and MMS. Supported formats are JPG, JPEG, and PNG — HEIC (the native iPhone format) isn’t supported yet, so tell customers to send a standard format if the upload fails. There’s also a size limit tied to carrier restrictions, so if an image won’t send, ask for a smaller file before you assume something’s broken.
There’s no extra charge for this. It runs on your normal bot usage, or gets folded into an unlimited bot plan the same way voice and email already do.

How do you build the prompt-based bot for image detection?
You build it in AI agents by choosing prompt based (not guided form) and writing the personality, goal, and info fields from scratch. Turn on autopilot, extend the response wait time to three or four seconds for image uploads, max out the message limit, and set the bot to sleep after manual takeover.
Start in AI agents and create a bot. Choose prompt based, not guided form — you want full control over the personality, goal, and info fields instead of a form-driven flow. Build it from scratch so you’re not stuck with settings from a template.
A few settings worth getting right from the start:
- Turn on autopilot so the bot can take over the conversation when triggered.
- Bump the response wait time from two seconds to three (or four) — image uploads take longer than text, and a bot that jumps in too fast will talk over the upload.
- Max out the message limit.
- Put the bot to sleep for a couple hours after you engage manually, so it doesn’t re-enter a conversation you’ve already taken over.
Writing the personality, goal, and info prompt
This is where the real work happens. The prompt has three parts inside GoHighLevel’s bot builder:
- Personality — how the bot talks and carries itself.
- Goal — the main intent. For an auto body example, that’s analyzing vehicle damage from photos, helping the customer understand what’s wrong, guiding them to the right repair path, and collecting what’s needed to book an estimate.
- Additional info — the image handling rules, the conversational guidelines, and anything specific to how you want damage described.
You have to explicitly tell the bot it can see and interpret images — it won’t assume that on its own. That instruction lives in the prompt, not buried in some setting.
Turn on appointment booking against a real calendar, and turn on human handover for every scenario available: when a customer asks for a human, when information is missing, and when the bot can’t resolve the issue. That last toggle matters most — the moment a customer gets frustrated, you want an easy exit to a real person.

How do you activate the bot through a chat widget automation?
You activate it by building a separate chat widget bot (not live chat, since only one live chat bot runs per account), triggering an automation on Customer Replied, filtering for that specific chat widget, and using the Update Conversation AI Bot Status action to turn the bot on only for that widget’s conversations.
This is the part people miss. You can only run one live chat bot per account, so if you already have a primary bot, you can’t just make this your live chat. Instead, build a chat widget bot — it works differently. A live chat bot is always on; a chat widget bot only turns on once triggered, which means you can run as many of them as you want, each tied to a different page or scenario.
Here’s the wiring:
- Create a funnel or page with a chat widget (not live chat) added to it.
- Build an automation triggered on Customer Replied.
- Add a filter for reply channel = chat widget, and specify which chat widget.
- Send an initial conversational message so the customer has something to reply to.
- Add a condition branch for when the customer replies.
- Inside that branch, use the Update Conversation AI Bot Status action to activate your image-detection bot specifically for that widget.
That last step ties it all together — the bot only turns on for people talking through that specific chat widget, not across your whole account.

What does it look like when a customer actually sends damage photos?
In the test scenario, a customer messaged in about being hit by another car, sent one photo, and the bot returned a full breakdown — affected areas, fender and headlight damage, whether glass was broken, and an estimated repair scope — then offered open slots to book an inspection, all from that single image.
The test scenario was a simple auto body shop bot. A customer messages in saying they got hit by another car, sends a photo, and the bot responds with a full breakdown: affected areas, visible damage to the fender and headlight area, whether there’s broken glass, an estimate of repair scope, and then it offers open appointment slots to book an inspection — all from one image.
“Do you guys see the dollar signs on this thing?”
That’s the moment this stops being a novelty. A bot that reads a damage photo and moves a lead straight toward a booked appointment is doing real work — it’s spotting the issue, guiding the conversation, and getting the customer to a booked slot without you touching it. That’s a different animal than a bot that just answers FAQs.

What’s the honest gotcha with this feature?
It’s brand new, and it shows. The first test image — an AI-generated damaged car photo — was too large to upload and threw errors. iPhone photos in HEIC format will fail outright until that format is supported. And because it’s new, there’s a real chance the bot gets stumped by an image it can’t parse well, so keep a human ready to step in, especially early on.
Also: this only works because the prompt is specific. A generic prompt gets you generic, half-useful responses. Budget real time — expect to spend more time refining the prompt than actually flipping the feature on.

How do you test and refine an image-detection bot?
Test with real images through the actual connected channel, not just the builder preview, checking that the bot’s description is accurate, its scoring language matches your prompt, and downstream automations like appointment booking fire correctly. Budget about an hour total: five minutes to enable image support, twenty for the prompt, fifteen for the workflow, and fifteen for testing.
Don’t ship this untested. Send real test images through the actual connected channel, not just the builder preview. Check that the bot’s description of the image is accurate, that its scoring language matches what you wrote in the prompt, and that downstream automations — like appointment booking — actually fire. Then run three to five different image scenarios to see how it handles variety, not just your best-case photo.
Rough time budget for the whole build: about five minutes to enable image support, twenty minutes to write a solid prompt, fifteen minutes to wire the workflow, and another fifteen minutes of testing. Call it an hour total — most of which should go into the prompt, since that’s what actually determines whether the responses are useful or generic.
Frequently asked questions
Does GoHighLevel’s AI image detection cost extra?
No. It runs on your existing bot usage or unlimited bot plan the same way voice and email messages already do — there’s no separate charge for image analysis itself.
What image formats does GoHighLevel’s bot support?
JPG, JPEG, and PNG are supported now. HEIC, the default format on iPhones, isn’t supported yet, so ask customers to switch formats if an upload fails.
Can I use image detection on my primary live chat bot?
Only if that live chat bot is your one primary bot for the account. If you already have a live chat bot in place, build the image feature into a separate chat widget bot instead, since GoHighLevel only allows one live chat bot per account.
Why isn’t my bot reacting to images correctly?
Usually it’s the prompt. The bot needs explicit instructions in the goal and additional info sections telling it that it can interpret images and how to describe what it sees — without that, responses come back generic or the bot ignores the photo entirely.
If you want the rest of the bot-building fundamentals before layering on image detection, start with the free GoHighLevel and AI training here, or work through the GoHighLevel Masterclass for the full step-by-step course. And if a question comes up mid-build, check the FAQ for common GoHighLevel questions.
