Owners of busy restaurants do not reply to their Google reviews. Not because they
don't care, but because it's forty minutes at eleven at night and it's the first thing
that goes. So Titanium drafts the replies for them, in the voice of that particular
business, and the owner reads them and sends.
Getting a model to write one convincing reply is an afternoon's work. Getting it to
write a hundred that don't read as though a machine wrote them is a different problem,
and it's the one worth explaining.
A reply goes through four passes. It gets drafted, then a second pass critiques that
draft and rewrites it, because a model told to avoid a phrase will satisfy the ban
lexically and leave the tell intact; you need something looking at the output rather
than the instruction. Then it's trimmed to a word budget that shifts with the star
rating, and finally checked against the rules that can't be broken. It is fenced off
from inventing facts: it cannot name a dish the reviewer never mentioned.
The part I'd point at, though, is the ledger. Google removes owner responses that
mirror each other across a profile, and a model has no memory of what it wrote
yesterday. So every reply that actually goes out is recorded, and the last thirty are
mined for repeated openings, closings and three-word runs, which are then handed to
the next draft as a list of things it may not say. Variation is enforced outside the
model, because it cannot be trusted to the model.
Scored against a fixed set of twenty-five reviews, the replies went from 3.5 out of 5
to 4.8. I run that set before and after every change to the prompt.
Without it, you are changing the wording and guessing whether it helped.
Most people building on these models are guessing.