How to Do Keyword Research with AI
AI has genuinely changed keyword research — but not the way the hype suggests. LLMs are superb at the language half of the job and structurally incapable of the data half. Here is a working method: where AI helps, where it hallucinates, a step-by-step workflow with copy-paste prompts, and the verification loop that keeps invented metrics out of your strategy.
Key Takeaways
Where AI wins
Language, not measurement
LLMs are excellent at expanding seeds, grouping intent, and mining questions — the parts of keyword research that are about how people phrase things
Where AI fails
Volumes and difficulty
An LLM has no live search data. Any volume, difficulty, or trend number it outputs is invented — verify every candidate against a real data source
The method
Generate wide, verify narrow
Use AI to produce a broad candidate list cheaply, then let real search data decide what survives into your strategy
Why Use AI for Keyword Research
Keyword research has always been two jobs wearing one name: a language job (imagining every way a shopper might phrase what you sell) and a data job (measuring which of those phrasings people actually search, and how hard they are to win). AI transforms the first job and cannot do the second. Understanding that split is the entire method.
It compresses the brainstorming phase from hours to minutes
The slowest part of manual keyword research is generating candidates: sitting with a blank sheet trying to imagine every way a shopper might describe your product. An LLM has read more product descriptions, forum threads, and reviews than any human researcher, so it produces a wide, plausible candidate list in one prompt. You still filter it — but filtering a list is far faster than inventing one.
It understands intent, not just strings
Traditional keyword tools group terms by shared words, which is why "leather bag repair" and "leather bag" land in the same bucket despite serving completely different searchers. An LLM reads meaning: it can separate buying intent from research intent, group fifty keywords into coherent topics, and flag the ones your store should not chase at all.
It scales to a whole catalog
Researching keywords for one product is a task; researching them for four hundred products is a program. AI is the only practical way to run the same disciplined process across every collection and product page — which matters, because the pages you never got around to researching are usually the ones quietly targeting nothing.
What it does not replace
AI does not replace search data, and it does not replace judgment. The volumes, difficulty scores, and trend lines that tell you whether a keyword is worth chasing come from tools that measure real searches. And the final call — which keyword each page owns — is a strategy decision about your store, not a text-generation problem.
Where AI Genuinely Helps
Five parts of keyword research are language problems, and LLMs are better at them than any manual process. These are the tasks worth handing over.
Seed expansion
Give an LLM ten seed terms and real context about your store, and it returns the long tail you would have missed: synonyms, attribute variants, use-case phrasings, gift framings, comparison queries. This is pure language work — exactly what the model is built for. The output is a candidate list, not a strategy; every term still gets verified against real data.
Intent clustering
Paste a flat list of keywords and ask the model to group them by what the searcher is actually trying to do — buy now, compare options, learn how, fix a problem. This mapping from keyword to intent to page type (product, collection, blog post) is the step most manual research skips, and it is the one that stops you writing a blog post for a buying query or a product page for a research query.
Question mining
Shoppers ask questions before they buy, and those questions are keywords: "is merino wool itchy," "how long do canvas sneakers last." An LLM generates the questions your category provokes — grounded in your actual products — faster than scraping People Also Ask by hand. These feed FAQ blocks, blog topics, and the answer-shaped content AI search engines quote.
Competitor-content analysis
Paste a competitor page — or the outlines of the current top results for a keyword — and ask what topics they cover that you do not, what searcher questions they leave unanswered, and what angle is missing from all of them. The model is fast at reading and comparing documents, which turns "study the competition" from an afternoon into a prompt.
SERP-pattern reading
The current search results page tells you what Google believes a keyword means: all product grids means buying intent, all listicles means research intent, a mix means the intent is split. Paste the top ten titles and ask the model to read the pattern — what page type dominates, what that implies about intent, and whether your planned page matches. It reads the pattern well; just remember it cannot fetch the SERP itself, so you supply what you see.
What AI Cannot Replace
Everything in this section is measurement, and an LLM measures nothing. It generates text that resembles what measurements look like — which is worse than admitting ignorance, because the output arrives formatted like a report.
Search volume
An LLM does not know how many people search for anything. It was trained on text, not on query logs. If you ask for volumes it will produce confident, plausible, fabricated numbers — and because they look exactly like real data, they are more dangerous than no data. Volumes come from tools with real clickstream or search-log data: Google Keyword Planner, Ahrefs, Semrush, or Search Console for queries you already rank for.
Keyword difficulty
Difficulty is a function of who currently ranks and how strong their pages are — live competitive data an LLM cannot see. It can reason about difficulty in the abstract ("branded head terms are harder than long-tail questions"), which is useful for prioritizing a raw list, but the actual score for an actual keyword needs a tool that crawls the actual SERP.
Trends and seasonality
Whether a term is growing, dying, or spiking every November is time-series data. A model with a training cutoff is structurally unable to tell you what happened to search behavior last month. Google Trends is free and answers exactly this question.
What you already rank for
Your own Search Console data — the queries where you show up, at what position, with what click-through — is the highest-signal keyword data you have, and no AI has access to it unless you hand it over. Starting AI research without it means re-deriving from guesswork what your store already knows from measurement.
The honest framing: asking an LLM for search volumes is like asking a well-read friend to estimate them. The guess will sound informed, and it will be a guess. The workflow below is built so that no guessed number ever reaches your strategy — AI proposes candidates, and only tools with real search data get to attach metrics to them.
The Workflow, Step by Step
Seven steps: ground the model in your real store, generate wide with AI, verify narrow with real data, and end with every keyword owned by exactly one page.
1
Assemble what you already know
Before any prompt: export your top queries from Google Search Console, list your best-selling products and collections, and write one sentence about who actually buys from you. This grounding is what separates useful AI output from generic output — a model prompted with your real catalog and real ranking queries expands from truth; a model prompted with "an online store" expands from cliché.
2
Expand your seeds with an LLM
Feed the model your seed terms plus the grounding from step 1 and ask for a wide candidate list organized by theme. Ask for how shoppers phrase things — problems, use cases, attribute variants, comparisons — and explicitly tell it NOT to output volume or difficulty numbers, so invented metrics never enter your sheet.
3
Cluster the list by intent
Paste the full candidate list back and have the model group it by searcher intent and map each cluster to a page type: buying-intent clusters to collection and product pages, research-intent clusters to guides and blog posts, question clusters to FAQ content. This produces the skeleton of a content plan, not just a keyword list.
4
Mine the questions
Run a dedicated pass for questions shoppers ask before buying in your category — pre-purchase doubts, comparison questions, care and durability questions. Sanity-check the output against real sources (People Also Ask, your own support inbox, product reviews): the model proposes, reality confirms.
5
Verify everything against real search data
This is the step that makes the workflow research instead of fiction. Take the surviving candidates into a real data source — Keyword Planner, Ahrefs, Semrush, or Search Console — and attach actual volumes and difficulty. Expect a meaningful share of AI-generated candidates to have no measurable volume at all; deleting them is the system working, not failing. What survives is a list that is both imaginative and true.
6
Read the SERP before you commit
For each keyword you plan to build a page around, look at what currently ranks. Paste the top titles into the model and ask it to read the pattern: what page type dominates, what intent that implies, whether your planned page matches, and what angle the current results miss. If the SERP is all product grids and you planned a blog post, the SERP wins — change the plan.
7
Assign one primary keyword per page
End with ownership, not a list. Every page gets exactly one primary keyword; no two pages get the same one. Skipping this step is how growing stores end up with three pages competing for the same term — keyword cannibalization that makes all three underperform. The clusters from step 3 make assignment mostly mechanical: one cluster, one owning page, siblings differentiated.
Where the last step leads: assigning one primary keyword per page is what prevents keyword cannibalization — the failure mode where your own pages compete against each other. If you are starting from ecommerce fundamentals rather than the AI workflow, the ecommerce keyword research guide covers the ground this one builds on.
Copy-Paste Prompts That Work
Each prompt does one job, carries real context, and forbids invented metrics. Replace the bracketed parts with your store's specifics — the more real detail you substitute, the less generic the output.
1
Seed expansion
I run an online store selling [products] to [audience — e.g.
"home cooks who buy premium kitchen tools as gifts"].
Our typical price range is [range].
Here are seed keywords from our catalog and Google Search
Console: [paste 10-20 terms].
Expand this into a list of 50 keyword candidates a real
shopper might type into Google. Include: attribute variants
(material, size, style), use cases, problems the products
solve, comparison phrasings, and gift-intent phrasings.
Organize by theme.
Do NOT include search volume, difficulty, or any metrics —
just the keywords.2
Intent clustering
Here is a list of keywords: [paste your candidate list].
Group them into clusters by searcher intent. For each
cluster, tell me:
1. The intent (ready to buy / comparing options / learning
how / solving a problem)
2. What kind of page should target it (product page,
collection page, buying guide, blog post, FAQ)
3. Which keyword in the cluster looks like the natural
primary, and which are supporting variants
Flag any keywords that don't fit our store and should be
dropped, and say why.3
Question mining
I sell [products] to [audience]. List 30 questions real
shoppers ask BEFORE buying products like these — doubts,
comparisons, sizing and fit worries, durability and care
questions, compatibility questions.
Phrase each one the way someone would actually type it into
a search box, not in polished marketing language. Group
them by buying stage: just researching / comparing /
about to buy.4
SERP-pattern reading
I'm deciding whether to target the keyword "[keyword]"
with a [page type you're planning].
Here are the titles of the current top 10 Google results:
[paste them].
Answer:
1. What page type dominates (product grids, listicles,
guides, forums)?
2. What does that imply about the searcher's intent?
3. Does my planned page type match — and if not, what
should I build instead?
4. What angle or question do the current results leave
uncovered?Why every prompt bans metrics: the moment a fabricated volume number lands in your spreadsheet next to real ones, you can no longer tell them apart. Keeping generation and measurement in separate tools is not a workaround — it is the design.
Common Mistakes
Every one of these comes from the same root error: treating the AI's output as finished research instead of raw material.
Trusting AI-generated metrics
The single most damaging mistake. If a volume or difficulty number came out of a chat window, it is fabricated — the model is pattern-matching what plausible numbers look like. Prompt the model to omit metrics entirely, and treat any that slip through as noise.
Prompting without context
"Give me keywords for a candle store" produces the same list for every candle store on earth — which means it produces keywords everyone is already chasing. The value of the workflow is proportional to the grounding you provide: your products, your customers, your ranking queries, your price point.
Skipping SERP verification
A keyword can have great volume, reasonable difficulty, and still be wrong for the page you planned — because Google has already decided the query deserves a different page type. Two minutes of looking at what actually ranks prevents weeks of building content the SERP will never reward.
Generating without assigning ownership
An AI-expanded keyword list is hundreds of terms. Spread them across pages without deciding which page owns which term and you have automated the creation of keyword cannibalization. The list is raw material; the strategy is the ownership map.
Running it once and never again
Search behavior moves, your catalog moves, and your Search Console data gets richer every month. Keyword research is a loop, not a project — which is the strongest argument for automating it rather than heroically redoing it by hand every quarter.
Automating the Loop for Shopify
Run manually, this workflow takes an afternoon per catalog section — and it decays, because search behavior and your catalog both keep moving. That makes it a natural candidate for automation, as long as the automation keeps the same honest division of labor.
Obsess AI’s keyword intelligence runs this guide’s loop for Shopify stores automatically, with the same structure: an AI strategy pass grounded in your actual store — your products, collections, store profile, and the queries you already rank for — produces the seed clusters and intent priorities; the seeds go to a real search-data provider for actual volumes and difficulty, so no model-invented metric ever enters the system; and every returned keyword is judged individually against your strategy, with off-audience and off-catalog terms dropped rather than padded into the list.
The output is the thing step 7 asks you to build by hand: an ownership map, where each page carries one primary keyword and the clusters stay stable as your catalog grows. Because the loop is automated, it re-runs as your store changes — new products get researched the day they exist, not the next time someone finds an afternoon.
If you are still choosing tooling, the best AI SEO tools roundup compares the landscape — general LLMs, traditional suites with AI features, and purpose-built systems — against exactly the generation-versus-verification split this guide is built on.
Frequently Asked Questions
Common questions about doing keyword research with AI.
Can AI do keyword research on its own?
Not end to end. An LLM handles the language half of keyword research extremely well — expanding seed terms, grouping keywords by intent, mining the questions shoppers ask — but it has no access to live search data, so it cannot tell you what people actually search for or how hard a term is to rank for. Complete keyword research is AI generation plus verification against a real data source such as Google Keyword Planner, Ahrefs, Semrush, or your own Search Console. Purpose-built systems close the loop by wiring both halves together.
Can ChatGPT give accurate search volumes?
No. ChatGPT and other LLMs are trained on text, not on search query logs, so any search volume, keyword difficulty, or CPC figure they output is a fabrication that happens to look like data. This is the most important thing to understand about AI keyword research: the numbers are confident, plausible, and invented. Use AI for generating and organizing keyword ideas, and get every metric from a tool with real search data.
How do I ask AI to do keyword research?
Give it context, a specific task, and a constraint against inventing data. Context: what you sell, who buys it, and your existing seed terms or Search Console queries. Task: one job per prompt — expand seeds, cluster by intent, generate questions, or read a SERP you paste in. Constraint: tell it explicitly not to output volume or difficulty numbers. Vague prompts like "give me SEO keywords for my store" produce the same generic list every competitor gets; grounded prompts produce candidates worth verifying.
What are the best AI tools for SEO keyword research?
Three categories cover the field. General LLMs (ChatGPT, Claude, Gemini) are the flexible option for expansion, clustering, and question mining — free or cheap, but blind to search data. Traditional keyword tools with AI features (Ahrefs, Semrush, Keyword Insights) bring real volumes and difficulty and increasingly add AI-assisted clustering on top. Purpose-built systems like Obsess AI’s keyword intelligence run the whole loop — AI strategy grounded in your store, real search-data verification, and per-page keyword ownership — automatically for Shopify stores.
Is AI keyword research reliable enough to build a content strategy on?
Yes, if you keep the division of labor honest: AI generates and organizes, real data verifies, and you decide. A workflow where every AI-suggested keyword is checked against actual search data before it enters the plan is at least as reliable as manual research — and dramatically broader, because the model surfaces phrasings a human brainstorm misses. A workflow that ships AI output unverified is not research; it is plausible fiction with a spreadsheet format.
How is AI keyword research different for a Shopify store?
Ecommerce keyword research is anchored to a catalog: most keywords should map to a product or collection page that can actually rank for them, and buying-intent terms matter more than raw volume. That makes grounding non-negotiable — the AI needs your products, collections, and existing rankings, not a generic description of your niche. It also raises the stakes on ownership: stores have many similar pages, so an unassigned keyword list turns into cannibalization faster than it would on a ten-page site.
More Keyword Resources
Aman Bedi, Founder, Obsess AI
Aman is the founder of Obsess AI and leads product and engineering on the Shopify-native AI content system. He works with Shopify merchants daily on keyword strategy, on-product SEO, blog content workflows, and the platform integrations that make all of it possible. The workflow in this guide is the same generate-verify-assign loop built into Obsess AI’s keyword intelligence, where it runs against real search data for Shopify stores automatically.
- AI Keyword Research
The Whole Loop, Run for You
Obsess AI’s keyword intelligence runs this exact workflow for your Shopify store — AI strategy grounded in your real catalog and rankings, every keyword verified against real search data, and one primary keyword assigned per page.
No credit card required