Most AI chatbot vendors are a black box. Yours will not be.
You cannot manage what you cannot see. A typical chatbot vendor hands you a widget and no way to tell what customers are actually asking, whether it answered them from your own content or quietly punted, or whether anyone came away satisfied. Helpforge hands you the report instead: what people ask, how well it answered, and exactly what to fix next.
Illustrative report below. The numbers are yours: once your agent is live, this fills in from your own visitors.
What a typical chatbot never tells you.
You paid to stop answering the same questions yourself, but most tools give you no way to check whether it is working. That is not a minor gap: it means you are trusting a system you cannot audit, with your customers on the other end of it.
The questions your customers actually type never reach you. Twenty different wordings of the same shipping question look like twenty unrelated rows in a raw log, if you can see the log at all, so the pattern that matters gets lost in the noise.
The traffic you would act on is invisible by default.A confident-sounding reply and a correct one look identical from the outside. Without a way to see whether an answer actually came from your content, or from the model filling a gap on its own, you have no way to know which replies you can stand behind.
Confidence is not the same thing as being right.Most widgets never ask the visitor how it went. With no thumbs up or down on a reply, an owner is left guessing at whether the bot is actually helping or quietly frustrating people who never bother to complain.
Silence is not the same thing as satisfaction.Even a vendor that shows you a log rarely tells you what to do about it. Finding the three questions your content does not cover, out of thousands of messages, is a research project most owners never have time to run, so the gap just stays open.
A pile of transcripts is not a plan.The report, and what each number in it means.
Every number below comes from your own visitors' real conversations, computed the same honest way throughout: a rate reads "not enough data yet" until there is something to measure, never a fabricated zero.
Every question is paired with the reply that followed it, then grouped by topic, not by exact wording. "Do you ship to Canada", "can I get this to Canada" and "Canada shipping?" collapse into one row with a phrasings count, so the list reads as your real question mix instead of a page of near-duplicates.
You see the pattern, not twenty copies of it.Every reply is classed as grounded (cited from your own content), uncited (answered but cited nothing), or unanswered (generation failed). The grounded share is the one honest proxy we have for "actually answered from your content" rather than filled in, and it is shown to you, not just asserted.
You can check the claim instead of taking our word for it.Every reply carries a thumbs up or down the visitor can leave. Satisfaction is up-votes over rated replies, kept as its own number because how a visitor felt and whether the answer was grounded can disagree, and a report that only showed one would be hiding the other.
Feelings and facts are two different rows, on purpose.The median server-side time to produce a reply, in seconds. Median rather than average, so one slow retry or a cold start does not move the number a single sluggish reply should not be allowed to move.
The typical wait, not the best case anyone cherry-picked.Every uncited or unanswered question, newest first, with the retrieval match it found (or "no reply" when generation failed). This is the worklist: the exact questions your content does not cover yet, not a research project you have to go run yourself.
Not a log to read. A list to work through.Seeing the gap is the easy half. Closing it is the other one.
A report that only shows you a problem is a complaint. Ours points straight at the fix: add a short answer for the exact question the bot could not cover, in your own dashboard, and it is grounded from that point on.
Before
- A gift-card-plus-promo question quietly goes unanswered, over and over
- You find out only when a customer complains, if you find out at all
- Fixing it means editing your whole site and waiting for a re-crawl
After
- The question sits at the top of "Needs attention" the next time you check
- You type a short curated answer straight into the dashboard, no re-crawl needed
- You can also turn a good answer already sitting in a transcript into a curated one directly, no retyping
One flat build fee, one flat monthly.
The report is not a paid add-on or a higher tier: it is what the dashboard looks like for every client, included in the monthly. No separate analytics fee, no seat count.
Free demo on your real site first. See your own agent, and the shape of its report, before you pay.
The insight report, answered honestly.
What does the insight report actually show me?
Your most-asked topics with a phrasings count, the share of replies actually answered from your own content versus not, a visitor satisfaction rate from thumbs up and down, the median time to produce a reply, and a "needs attention" worklist of every question the bot could not answer well, newest first. It is the dashboard's first screen, not a separate paid tier.
What counts as "answered from your content" versus "uncited"?
A reply is grounded when it cites at least one of your own pages, which is the closest honest proxy we have for "actually answered from your content" rather than filled in. Uncited covers both small talk and a real content gap, and the match score shown next to it in "needs attention" tells you which one you are looking at, so you are not left guessing.
How does topic clustering work, and could it merge two different questions by mistake?
Questions are grouped by their shared meaningful words once filler words and word order are set aside, so reorderings and near-duplicates collapse into one topic with a phrasings count. It is deliberately conservative: the merge only happens on strong overlap, tuned so "ship to Norway" and "ship to Japan" stay apart while five wordings of "what is your return policy" become one row. A wrong merge would misreport your traffic, a missed one only under-groups it, so it is set on the safe side.
What is the satisfaction rate based on?
A visitor's own thumbs up or down on a specific reply, nothing inferred. Satisfaction is up-votes divided by everything rated, and it reads "no ratings yet" rather than a fabricated 0% until at least one visitor has actually rated something. It is tracked separately from the grounded rate on purpose: a correct answer can still get a thumbs down, and a report that hid that gap would not be honest.
How is response time measured?
The server-side time to generate each reply, in milliseconds, reported as the median across every timed reply in the window. Median rather than average so a single slow retry or a cold start cannot drag the number a typical visitor never actually experienced.
What do I do when the report flags something the bot could not answer?
Add a short curated answer for that exact question in your dashboard. It takes effect immediately, no re-crawl required, and the next visitor who asks it gets a grounded reply. You can also promote a good answer straight out of a real transcript into a curated one, pre-filled, rather than retyping it from scratch.
See your own report take shape.
Give us your website URL and we build a working demo of your agent, usually within 2 business days. You get a private link to test it and a look at the report screen your real report will use, before any commitment. No call, no credit card.