Malaysian agencies bypass the default Google Keyword Planner workflow by feeding GSC query exports and Shopee/Lazada autocomplete terms through LLM clustering APIs, normalising Manglish and Bahasa Malaysia variants into intent-classified JSON graphs for programmatic SEO.
The Klang Valley Query Mix Problem
The default keyword research flow — log in to Ahrefs, export rows, merge in Excel — collapses in Malaysia. Search demand here is trilingual in practice. A user in Petaling Jaya types “servis aircond terdekat”; the same user searches “aircon repair near me” a week later. Ahrefs reports those as separate rows with separate keyword-difficulty (KD) scores. A manual workflow will treat them as separate topics, and the agency will later publish two pages that compete for the same intent.
Malaysian agencies now treat the surface text as noise. They normalize all variants of the same intent into one canonical cluster, and the normalization job is done by an LLM over a corpus of 5,000+ query rows. The prompt asks the model to collapse morphological differences specific to Bahasa Malaysia: reduplication (baju-baju), the meN- verb prefix (mencari, memilih), and Manglish length markers (“dekat sini”, “area sini”, “sini”). The output is a lookup table where the cluster ID — not the surface string — drives every downstream decision.
Pulling Low-Volume Queries from GSC
Google Keyword Planner in Malaysia masks anything below roughly 1,000 monthly searches for the MY region. That cutoff is useless for a market where high-intent queries like “berapa harga servis aircond” live in the 50–500 search range.
Agencies therefore pull the raw string list from the Google Search Console query report (a 16-month export, impressions filter at 10). GSC shows the exact low-volume strings users typed to reach a page that already ranks. To cover demand that does not yet map to an existing page, agencies add a second pull from the DataForSEO API with location code 2384 (Kuala Lumpur) and a 50-searches-per-month floor. The two lists are unioned and sent, in batches, to the LLM for diacritic, spacing, and casing normalization (e.g., “anak ayam” vs “anakayam”, “chicken rice pj” vs “chicken rice petaling jaya”). Only then does the cluster count begin.
Intent Tagging via LLM API Calls
Once clustered, every node needs an intent tag so the page template can be chosen correctly. Malaysian agencies use four buckets: transactional, informational, navigational, and geolocation. The geolocation bucket is non-negotiable in the Klang Valley because queries like “water chiller supplier selangor” or “kedai makan halal puchong” carry a physical radius, not just a product intent.
The prompt feeds 50 keywords per API call and forces a JSON schema response:
{“cluster”: “water-chiller-supplier”, “intent”: “geolocation”, “geo”: “selangor”, “page_template”: “/supplier/{geo}”}
Manual tagging of 2,000 rows takes roughly two business days. The batch Gemini Flash or GPT-4o run finishes in under 15 minutes. Agencies then validate a 5% sample against human judgment before the tags move into the content plan.
Shopee and Lazada Autocomplete as Keyword Input
For product-led clients, Google is not the only search surface with reliable demand data. Shopee and Lazada autocomplete capture the vocabulary Malaysian buyers actually type before clicking “buy”. A base term like “baju kurung” returns 15 suggestions on Shopee, including “baju kurung moden online”, “baju kurung shanghai”, and “baju kurung kanak-kanak” — strings that barely register in Google volume data.
Agencies scrape the marketplaces’ public autocomplete endpoints, union the output with the GSC list, and push the combined set through the same LLM clustering step. The practical result: a keyword with strong Shopee volume and near-zero Google volume still gets assigned a page template and a listing-title instruction. For Chinese-speaking segments, Lazada’s Mandarin autocomplete (e.g., “保温杯” for thermal flasks) adds a third-language node that Google Keyword Planner ignores entirely.
Structuring the Keyword Graph for Writers
The final deliverable is not a flat spreadsheet. Agencies output a keyword graph — a JSON tree in which every node holds the canonical keyword, the summed monthly volume across all variants, the intent tag, the geo tag, and the target page template. Writers consume this graph through Notion or Airtable as a filtered “content brief” view: article title, target cluster, linked internal pages, and the exact long-tail variants to mention in body copy.
The graph also prevents self-competition. If two proposed landing pages map to the same canonical cluster, the duplication is caught at planning stage rather than after publishing. For Klang Valley clients, the graph is typically sorted by geo node — Puchong, Petaling Jaya, Shah Alam — so agencies can confirm nobody is building three identical pages for the same service within the same municipality.
| Keyword source | Tool / method used | Reason for inclusion |
|---|---|---|
| Google Search Console (16-month export) | CSV export + LLM variant normalization | Captures low-volume queries Keyword Planner hides |
| DataForSEO API (location 2384, KL) | Python pull, 50 searches/month floor | Fills long-tail demand absent from existing pages |
| Ahrefs Keyword Explorer | KD filter below 30 | Picks cluster nodes with weak SERP competition |
| Shopee / Lazada search autocomplete | Marketplace endpoint scrape | Adds transactional product terms missing from Google |
| OpenAI GPT-4o / Gemini Flash | Batch JSON intent prompts | Clusters variants and tags intent at 50 keywords per call |
Ready to Accelerate Your Digital Growth Strategy?
Partner with an industry-leading digital agency to upscale your infrastructure today.







