All posts

Mastering Audience Research With AI: A Step-by-Step Guide

A practical two-pass method for audience research with AI: read the content your market rewards to find your angle, then profile the people behind that engagement on the Big Five to find the register they trust. Ends in copy you can ship.

You did the audience research. You have a persona: Marketing Mary, 34, marketing manager at a mid-size SaaS company, cares about ROI and hates busywork. You pinned it to the wall. And your copy still lands flat.

You are not the only one who has noticed. Corey Haines' customer research skill lists persona anti-patterns, and the first one is "Don't name them cutely ("Marketing Mary")", on the grounds that it is usually a distraction. The same page insists that "Personas should be built from research, not invented."

The persona is not wrong, then. It is answering a different question than the one that matters. It tells you who your audience is. It does not tell you what to say to them. Those are two separate research jobs, and almost everyone stops after the first.

This guide does the second job. It is a two-pass method, both passes AI-assisted, and it ends where audience research is supposed to end: in a sentence you would actually publish.

What audience research actually is#

Audience research has two layers, and the gap between them is the whole problem.

Demographics are the who: age, job title, company size, location. They are easy to collect and they tell you who to reach. Psychographics are the wiring: what a person values, fears and aspires to, and how they are disposed to think. They tell you what will move that person once you have reached them.

A quick word on which matters more, because the internet will tell you psychographics simply beat demographics. The research does not support that. A study of more than 45,000 people by Sandy, Gosling and Durant (2013) found both explain only small amounts of behavior, and which one wins depends on the behavior. The honest position is that psychographics add predictive power on top of demographics, and the two work best together. Demographics point you at the room. Psychographics tell you what to say once you are in it.

Getting the second part right is not a nicety. McKinsey's research puts personalization done well at a 5 to 15 percent lift in revenue, and its 2021 personalization report found that faster-growing companies drive 40 percent more of their revenue from personalization than slower-growing ones. The money is in the second layer. Most research stops at the first.

Pass 1: read the content your audience rewards#

Your audience is already gathered somewhere, engaging with content every day. What they engage with is a record of what they care about. You do not have to survey them. You have to read what they have already voted on with their attention. This pass needs no special tool and it ends with a drafted headline.

Step 1. Find the room#

Be concrete about where your market actually talks. Do not stop at "they're on LinkedIn." Produce a short, specific list:

  • The three or four subreddits, forums or communities where your category is discussed. Search your category on Reddit, sort by top of the past year, and note which subreddits the winning posts live in.
  • The five to ten creators or competitor accounts your buyers follow and comment under.
  • The hashtags or search terms under which the conversation happens.

The output of Step 1 is a list of maybe a dozen places, not a feeling. If you cannot name them, that is your first finding: you do not yet know where your audience is, and everything downstream is guesswork until you do.

Step 2. Rank what works, with AI#

Now read the content that earned engagement, at volume, and let a model do the sorting. Collect the top 20 to 30 posts from the rooms in Step 1, the ones with the most comments and shares, and copy each post together with its top comments into a single document. Then paste it into any capable LLM with this:

Promptcopy this into ChatGPT, Claude, or any LLM
Here are 25 high-engagement posts and their top comments from communities
where my target audience gathers. Do four things:

1. Cluster the recurring PAIN POINTS people express, ranked by how often
   they appear. Quote a representative line for each.

2. Cluster the TOPICS people show up for, ranked by engagement.

3. Tag each post's HOOK TYPE (contrarian take, personal failure story,
   how-to, data reveal, hot take) and tell me which hook types win most.

4. List the exact PHRASES and vocabulary this audience uses for its
   problems, so I can mirror their language rather than my own.

You are reading three signals at once: the topics they show up for, the pains they state in the comments, and the hook types that make them react. That third one matters more than it looks. Berger and Milkman, analyzing which New York Times articles got shared most, found that "content that evokes high-arousal positive (awe) or negative (anger or anxiety) emotions is more viral," while low-arousal emotions like sadness travel poorly. So when a particular hook keeps winning in your market, it is telling you which emotional register your audience acts on.

The output of Pass 1: an angle and a first line#

What an audience amplifies is a record of what it cares about. It reveals the audience's values, fears and wants, before you know a single name. A pain quoted in thread after thread is not an anecdote; it is a market. A hook that keeps winning is not a coincidence; it is the emotional door your audience opens.

So end this pass with something you can use. Here is what that looks like worked through. Say the LLM comes back with a top pain of "nobody can prove our content actually drives pipeline" and a top-winning hook type of the personal-failure confession. You draft:

Angle: stop defending content with traffic numbers nobody upstream believes; show the pipeline it touched.

Opening line: "We published 140 posts last quarter and I could not tell you which one closed a deal. Here is the tracking I wish I had set up on day one."

That opening line is the first concrete product of the method, built from a real pain in the audience's own words, and you got it for free. What this pass cannot give you is who these people are underneath the content. For that you need the second pass.

Pass 2: profile the people behind the engagement#

Here is the catch Pass 1 cannot solve. The same post is engaged with by very different people, and content analysis alone treats them as one blur. Two people upvote the same thread: one is a cautious, detail-driven operations lead, the other an impulsive, novelty-seeking founder. The same content reached both. The message that converts them is not the same message. To see that difference you have to profile the people, not the posts.

Step 3. Collect the people, and what they reshared#

Widen the collection from posts to the accounts behind the engagement: everyone who commented, liked or, most tellingly, reshared. A reshare is a stronger signal than a like, because people signal identity by what they choose to repost to their own audience. Collecting this by hand across several networks is the point where the method stops being a spare-afternoon job. This is the step buzzabout automates, pulling the engaged profiles across Reddit, TikTok, YouTube, X, Instagram and LinkedIn in one pass. It is worth naming the category once: this is social media research, not social listening. Social listening counts how often a keyword appears. Here you are reading the people.

Step 4. Score each person on the Big Five#

Personality is where "what they care about" becomes "how they are wired." The most widely used model in personality science is the Big Five, or OCEAN: Openness (appetite for new ideas and experiences), Conscientiousness (discipline and care), Extraversion (sociability and drive), Agreeableness (warmth and cooperation), and Neuroticism (sensitivity to worry and negative emotion).

You cannot interview thousands of engaged accounts to score them. You do not have to, and this is the part that surprises people: personality is predictable from public digital behavior. Kosinski and colleagues showed in 2013 that Facebook Likes alone predict a person's openness nearly as well as having them sit the personality questionnaire itself. Two years later, the same group found that a model beats the people who know you, and published exactly how much evidence it needs to do it.

10 Work colleague 70 Friend 150 Family member 300 Spouse
Likes a model needs to judge your personality better than a human who knows you

Three hundred Likes and the model reads you better than your spouse does. The content a person engages with is a validated signal of who they are, which is the entire premise of this pass. buzzabout's profiling scores each engaged account on the five factors from exactly this kind of public footprint.

Michal Kosinski, who ran that research, explains the mechanism and its consequences here:

One honest limit, stated plainly: scores inferred from public behavior are estimates, not administered clinical tests, and they only cover people who engage in public. Treat the output as a strong signal, not a diagnosis.

Step 5. Cluster by temperament, not by age#

Now group the scored people by psychographic similarity and read the clusters. This is where profiling earns its keep, because the clusters do not line up with demographics.

To show what the output looks like, here is one real run of this method on a marketing audience: 300 people who engaged in a year of audience-research conversation, each scored on the five factors. Sorting them by temperament produced four clean segments.

Segment C, methodical 105 · 35% Segment A, outgoing 91 · 31% Segment B, anxious 73 · 25% Segment D, wry 27 · 9%
One audience, sorted by temperament rather than demographics (296 profiles)
Segment People Mean age The tone they write in
A: outgoing, unruffled 91 34.7 authoritative, conversational
B: anxious, high-strain 73 29.0 critical, analytical
C: methodical, reserved 105 37.0 professional, authoritative
D: wry, casual 27 30.7 wry, casual

Look at the age column. The four groups sit between 29 and 37, effectively the same age. No demographic filter separates them. Their writing registers are not remotely the same.

That is not an accident of this one sample. Across these 300 people, age barely tracked personality at all. It explained less than one percent of the variation in openness. Even its strongest link, to conscientiousness, was modest.

0.7 Openness 1 Agreeableness 1.4 Extraversion 6.5 Neuroticism 16.5 Conscientiousness
How much of each personality trait is explained by age (share of variance, %)

Read that chart as the reason the Marketing Mary persona failed. It fixes an age and quietly infers a personality from it. The data says age tells you almost nothing about how a person is wired. Sort your audience by temperament and you get groups a demographic filter can never see.

The output of Pass 2 is this segment map: your audience, grouped by how its members are disposed to think, each group with a personality signature you can write to. To be clear about what this proves: the segment table shows you what the method produces. It does not, by itself, prove that writing to those segments converts. That proof comes next, and it does not come from us.

Pass 3: turn personality into words#

A segment map is not the deliverable either. The deliverable is different copy. This pass turns each temperament into the specific persuasion move it responds to.

Why this is worth doing at all#

Start with the evidence that it works, because it is not ours. Matz, Kosinski, Nave and Stillwell ran three real advertising field experiments reaching over 3.5 million people, matching ad copy to personality. Their result, verbatim: appeals matched to people's "extraversion or openness-to-experience level resulted in up to 40% more clicks and up to 50% more purchases than their mismatching or unpersonalized counterparts."

More purchases 50 More clicks 40
Personality-matched ad copy vs mismatched or generic, maximum uplift (Matz et al., PNAS 2017)

Same product, same audience, same spend. The only variable was whether the words fit the personality. Sandra Matz, who led that work, walks through what the experiments actually did:

Step 6. Match the technique to the temperament#

Now you need named techniques rather than intuition. Corey Haines' marketing-psychology skill is a good working catalogue: 72 named models, each with a definition and a marketing application, in a single open-source file.

One thing to be straight about before using it. That library organizes its techniques by marketing challenge, and it never mentions the Big Five. So the pairings below are our reasoning, not Haines'. What comes from the library is each technique's definition and application, quoted directly. Treat the pairings as strong starting hypotheses to test on your own audience, in exactly the spirit of the library's own warning not to invent details about an audience you have not researched.

Your segment scores high on Openness

→ First Principles, Inversion, the Pratfall Effect

Curious, novelty-seeking readers enjoy reasoning a claim out rather than being handed a conclusion. The library describes Inversion as asking "What would guarantee failure?" instead of "How do I succeed?", and the Pratfall Effect as the fact that "Competent people become more likable when they show a small flaw."

Your segment scores high on Conscientiousness

→ Authority, Commitment and Consistency, the Goal-Gradient Effect

Deliberate buyers do the reading, so credentials and documentation reward the effort they were going to make anyway. Get the small commitment first: the library advises to "Get small commitments first (email signup, free trial)."

Your segment scores high on Extraversion

→ Mimetic Desire, Social Proof, the Unity Principle

Sociable, status-attentive readers treat consensus as a positive signal. Mimetic Desire is the sharpest tool here: "People want things because others want them. Desire is socially contagious."

Your segment scores high on Agreeableness

→ Reciprocity, Foot-in-the-Door, Liking and Similarity

Warmth persuades the cooperative more than proof does. The library's application line for Reciprocity is the whole playbook: "Give value before asking for anything."

Your segment scores high on Neuroticism

→ Regret Aversion, Status-Quo Bias, Activation Energy

Threat-sensitive readers stall rather than object. Remove the risk and the friction: "Money-back guarantees, free trials, and "no commitment" messaging reduce regret fear."

What backfires, which is the more useful half#

Knowing which lever to pull matters less than knowing which one will blow up in your face. The same technique that converts one temperament repels another, and this is where most personalization advice stops short.

If your segment is high in Reach for Avoid, because it backfires
Openness First Principles, Inversion, Pratfall Effect Social proof, mere exposure, "the industry standard since 1994"
Conscientiousness Authority, proof, Commitment and Consistency Countdown timers and manufactured urgency
Extraversion Mimetic Desire, Social Proof, Unity Long technical proofs that delay the payoff
Agreeableness Reciprocity, Foot-in-the-Door, Liking Door-in-the-Face, the deliberate over-ask
Neuroticism Regret Aversion, guarantees, Status-Quo Bias Scarcity stacked on loss aversion

The Openness row is the one most marketers get wrong. Social proof is the reflex move in almost every campaign, and on a novelty-seeking, unconventional audience it is actively counterproductive: they are the people least moved by what everyone else is doing. The Conscientiousness row is the second. A countdown timer aimed at a careful, future-oriented buyer reads as pressure, and they discount it. The library hedges its own scarcity entry for the same reason, with three words: "Only use when genuine."

Step 7. Write the claim twice, then test it#

This is the step that turns everything above into copy. Take one claim about your product and rewrite it for two of your segments, using the levers from Step 6.

To make this real rather than theoretical, we took two of the segments from the run above and asked buzzabout's assistant one question about each: what makes this group distrust a marketing claim. It answered by pulling that segment's own comments back out. (This is an illustration of the output, not a second piece of proof. The proof that matched copy converts is Matz's experiment.)

Segment C, the methodical, reserved group, distrusts anything polished before it is proven:

"your value proposition says nothing a competitor couldn't also claim."

Don Schuerman, LinkedIn (21 likes, 3 comments)

Segment D, the wry, casual group, distrusts spin, and trusts only what it can observe:

"Hyper-targeting is a myth. A very expensive myth."

Peter Weinberg, LinkedIn (126 engagements)

Now take a single claim, say a project management tool that "helps teams move faster," and write it for each. For segment C, lead with authority and specifics, the levers a conscientious reader rewards:

The 200 teams with the highest retention on our platform all standardized the same three workflows in week one. Here are the before-and-after cycle times.

For segment D, drop every ounce of hype and admit the limit, which is the Pratfall Effect doing its work on an audience that is scanning for spin:

Most tools promise to transform how your team works. This one makes your status update take thirty seconds instead of your whole standup. That is the pitch.

Same product, same fact, two registers, each matched to a temperament. Do not take our word that the difference matters. Put both versions in front of the right segments and measure. The end of this method is an A/B test on your own list, not a claim.

Where to stop: the ethics of this#

This is powerful enough to misuse, so draw the line explicitly rather than leaving it implied.

The clearest statement of it in the skill library is in its prospecting skill, which instructs: "Never target or infer sensitive traits. Don't qualify, segment, or personalize on health, financial hardship, political belief, sexuality, religion, or other protected/sensitive attributes". Personality traits are not on that list, so trait-based segmentation is not the thing that guardrail forbids. But one row of the table above sits uncomfortably close to it.

That row is Neuroticism. Inferring who in your audience is most anxious, and then deliberately stacking urgency and loss language on those specific people, is amplifying anxiety in the readers least able to discount it. The technique works. That is the problem with it. Use the Neuroticism row to remove risk, which is what guarantees, free trials and one-click migration do, and not to manufacture fear.

The general test is simple: matching a message to a personality means speaking to someone in the register they actually trust. It does not mean fabricating claims. If you would be uncomfortable explaining the targeting to the person you used it on, that is the signal.

What to do this week#

The whole method, as a checklist:

  1. List the rooms. Name the specific communities, creators and hashtags where your market engages. If you cannot, that is finding number one.
  2. Rank what wins, with AI. Paste 25 high-engagement posts and their comments into an LLM. Get back ranked pains, winning hook types, and the audience's own vocabulary.
  3. Draft your angle. Turn the top pain and top hook into an opening line, in their words. That is Pass 1 done, at zero cost.
  4. Collect the people. Widen from posts to the accounts that engaged, and what they reshared.
  5. Score and cluster. Profile the engaged accounts on the Big Five and group them by temperament, not age.
  6. Match technique to temperament. Use the table above, and pay more attention to the "avoid" column than the "reach for" one.
  7. Write twice, then test. Rewrite one claim for two segments and A/B it on your own audience.

Steps 1 to 3 you can do this afternoon. Steps 4 and 5 are where a tool like buzzabout does the collection and scoring you cannot do by hand. But the shape of the work is the same at every scale: read what your market rewards to find your angle, read the people behind it to find their register, and do not call the research finished until it has changed a sentence.

Because that is the point a persona never reaches. "Marketing Mary, 34, cares about ROI" names your audience. It cannot write your copy. What can write your copy is knowing what your market keeps asking for, and which of the four people hiding inside that one persona you are talking to right now.

Frequently asked questions#

What is audience research? The work of understanding a market well enough to reach it and move it. It has two layers: demographics (who they are: age, role, location) and psychographics (how they are wired: values, fears, personality). Demographics tell you who to reach; psychographics tell you what to say. Most research stops at the first layer, which is why so many well-targeted campaigns still miss.

How is AI used in audience research? Two ways in this method. First, a general LLM clusters the pains, topics and hook types out of content your audience already engages with, which used to take a strategist a week. Second, a model scores the personality of the people who engage, from their public behavior, which cannot be done by hand at any useful scale. Research going back to 2013 shows personality is predictable from digital footprints accurately enough that a model beats a person's own friends and family.

Are demographics or psychographics more important? Neither alone. A study of over 45,000 people (Sandy, Gosling and Durant, 2013) found both explain only modest amounts of behavior and which wins depends on the behavior, and recommended combining them. Use demographics to decide who to reach and psychographics to decide what to say once you do.

How do you use psychographics to personalize marketing? Segment your audience by personality rather than age, then match each segment to the persuasion technique its temperament responds to, and rewrite the same claim once per segment. A high-conscientiousness segment rewards authority and proof; a high-openness segment rewards a fresh reframe and an admitted flaw; a high-neuroticism segment rewards removed risk. Just as important is what to avoid: social proof underperforms on high-openness readers, and urgency underperforms on high-conscientiousness ones.

Can I do audience research without a special tool? Pass 1 entirely, yes: finding the rooms, ranking the winning content with an LLM, and drafting your angle needs nothing you do not already have. Pass 2, profiling the individual people who engage and scoring them on personality across several networks, is the part that needs collection and scoring at scale.

What is the Big Five (OCEAN) model? The most widely used framework in personality science, describing five traits: Openness, Conscientiousness, Extraversion, Agreeableness and Neuroticism. It matters for marketing because personality predicts which kind of persuasion a person trusts, and because those traits can be estimated from the public content a person engages with.

Note on the worked example. The 300-profile segment run used to illustrate Pass 2 and Pass 3 is drawn from twelve months of audience-research conversation across Reddit, LinkedIn, X and YouTube, 307 posts and 1,594 comments, with 300 engaged accounts scored on the OCEAN model and age-trait correlations computed on the 231 profiles with a resolved age. Those figures illustrate what the method produces on one audience. The evidence that matching copy to personality lifts response is Matz et al., PNAS 2017, an independent field experiment, not our data. Technique definitions are quoted from Corey Haines' marketing-skills library, which is MIT licensed; the pairings of technique to personality trait are ours, not the library's. Personality scores inferred from public behavior are estimates, not administered tests.

See what the people behind your market's conversations are actually made of. Start at buzzabout.ai.