Aug 17, 2026

CrowdinAI Visibility Audit

crowdin.com · localization management platform (TMS) · 19,743 URLs in sitemap (49 core, 349 blog, 19,345 UGC project pages) · 21 language subdomains
Prepared by ESA Digital · Screaming Frog crawl (898 URLs) · Google Keyword Planner · DataForSEO Labs & LLM Mentions · CrUX & PageSpeed Insights · Common Crawl · live bot checks · 12 ChatGPT spot-checks

1.Executive summary

The foundation is sound, and that is the right place to start. robots.txt withholds nothing: rule-to-URL matching across all three sitemaps puts 0 of 19,743 URLs under a disallow. LLM crawlers are served the same site a human is — 13 AI user agents across 9–10 URLs, repeated over 3 runs, returned responses byte-identical to a browser baseline, with zero variance between runs and no CDN challenge in 130+ requests. JavaScript is not a barrier either: on none of the 888 matched pages does rendering add meaningful text. Editorially the company is strong — 311 of 311 blog articles (100%) carry a named author, there are 39 customer stories with a measurable result in 38 of them, and the blog's median E-E-A-T score is 87. Crowdin can produce content that models trust. The problem is that almost none of that strength reaches the pages where buying decisions are made.

Everything we found collapses into three groups. First, 98% of the site is a templated UGC catalogue: 19,345 /project/* pages sit in the index and the sitemap, are statistically identical to one another (5-gram Jaccard median 0.993–0.995 on two independent samples, median 1 unique token per page, 562 words of interface chrome, 0 H1, 0 structured data) and are 100% orphaned — the /projects listing emits zero links to them in raw HTML. Second, the money pages do not give a model facts to work with: query fan-out coverage is 53%, the price appears nowhere in text (0 occurrences of a dollar figure in the entire /pricing HTML — only in JSON-LD), the main comparison table conveys 49% of its cells through icons a text pipeline cannot read, and 44 of 46 commercial pages score below 60 on E-E-A-T against a blog median of 87. Third, the authority signals exist but are not published: Wikipedia, Wikidata, a 507-follower GitHub organisation, G2 with 662 reviews, Capterra with 175 — and not one of them is declared in sameAs on any page but the homepage.

The strategic finding is the one worth acting on first. Across 12 spot-check prompts, ChatGPT mentioned Crowdin in 11, at an average position of 2.9 — but named it best in none, and never once cited crowdin.com as a source. Measured domain citation counts put Crowdin at 15 against Smartling's 203, Smartcat's 134, Weglot's 117, Phrase's 107 and Lokalise's 58 — ninth of ten — while Crowdin's own AI search volume, 643, is higher than Lokalise's 441. Demand for the brand exists; the material a model would quote does not. The model is not distorting the site. It is summarising it accurately: asked to characterise Crowdin it reliably returns "open source, community, developer-heavy", which is exactly what the pages say — the homepage H2 reads "Built by developers, for developers", /pricing carries a "Crowdin for Open Source" section, and the phrases "for product teams" and "for SaaS companies" appear on none of the ten money pages.

That symmetry is what makes this tractable. Because the gap is in what the pages state rather than in whether machines can reach them, most of the highest-impact work is publishing and templating, not rebuilding: one directive change retires 19,345 thin URLs, one Astro partial edit adds grounded entity data to 358 pages, one pull request puts Crowdin into a 408-star awesome-list it is currently absent from, and one component already working correctly on a neighbouring page fixes the comparison table. Ten of the 73 checks are Blocked rather than passed, across six areas — CTR gap and pages-decay, LLM traffic share and ARPU, quantitative share of voice, Perplexity and Claude coverage, Reddit placements and contractor pitch targets. Those are marked as not measured, never as zero, and Appendix B lists exactly what to send us to close each one.

Overall score
57
/ 100
Needs work
SectionScoreWeightComment
AI crawler access86×2 7 PASS; one High (forum throttles LLM bots), one Medium (no price in text)
AI answer readiness46×2 fan-out 53%, extractability 65.1, 44 of 46 money pages E-E-A-T below 60
Semantics & content62×1.5 297-keyword gap; 89 keywords sit on pages that already exist
Technical42×1.5 3 Critical, 3 High, 4 Medium (crawler's own health score)
Visibility in models30×1 15 citations vs Smartling 203; absent from the category brand dataset
E-E-A-T64×1 authorship 100%, but no external profiles in sameAs
Structured data52×1 zero syntax errors, two policy risks
External placements58×1 2,026-donor gap; awesome-i18n absent; Trustpilot 3.8/18
Why 57 and not lower. Our scoring model caps the overall score at 39 when AI crawlers are blocked — nothing else matters while models cannot read the site. That cap does not apply here: robots.txt disallows 0 of 19,743 URLs, LLM bots receive byte-identical responses to a browser, and JavaScript hides nothing. Crowdin's problem is content and publishing, not access, which is a substantially better position to start from and the reason the roadmap is mostly editorial rather than infrastructural.

Why "Visibility in models" is 30 and not a share-of-voice figure. The 30 reflects measured citation counts from the LLM Mentions dataset (Crowdin 15 vs Smartling 203) plus absence from the 42-brand category dataset. It is not a share of voice: quantitative SOV requires Ahrefs Brand Radar, which was outside this run, and we do not substitute a different method for it. The 12 prompt spot-checks describe 12 specific answers and must not be read as a 92% visibility rate — on a previous engagement spot-checks implied roughly 11% share against an actual 0%.

Checks summary

#CheckStatusComment / action
AI access
1robots.txt coverage PASS 0 of 19,743 sitemap URLs under any disallow; verified by rule-to-URL matching, not by reading the rules
2Rules for named AI bots PASS No Disallow: / for GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot, Google-Extended, CCBot on any host
3Live bot check (UA discrimination) PASS 13 UAs × 9–10 URLs × 3 runs — byte-identical to browser baseline; 63/63 stability measurements identical
4CDN / WAF blocking PASS* No Cloudflare/Akamai signals in 130+ requests; assessed from headers only — not confirmed by the client
5Forum bot rate limit FIXhigh community.crowdin.com returns 429 to GPTBot and ClaudeBot (gptbot_crawler_rate_limit); ceiling ~1 request / 30 s. One Discourse setting
6Common Crawl presence PASS apex 1,225 records, blog 131, docs 426, forum 73 (CC-MAIN-2026-30); Wikipedia control 72,852
7Content without JavaScript PASS 0 of 888 pages gain >20% text from rendering, across 5 platforms (Astro, Next.js, Starlight, Discourse, legacy app)
8Dead legacy in robots.txt FIXlow 6 of 10 rules point to URLs that return 404 — harmless today, a trap at the next URL refactor
9Price legible to a text reader FIXmedium 0 dollar figures in the entire /pricing HTML, raw or rendered — price exists only in JSON-LD
Technical
10Response codes (sitemap URLs) PASS 898 crawled URLs, 100% return 200; zero 4xx/5xx among sitemap pages
11Custom 404 PASS True 404 code with a branded page — not a soft-404
12HTTPS, host canonicalisation, HSTS PASS http→https and www→non-www 301; HSTS max-age=31536000; includeSubdomains
13Duplicate titles PASS 0 of 898 — every title unique
14Duplicate and multiple H1 PASS 0 duplicate H1 and 0 pages with two H1s
15Canonical correctness PASS 0 canonical conflicts, 0 missing on core — no canonical pointing at a redirect or 404
16Core Web Vitals — mobile FIXcritical Origin field LCP 6,815 ms, INP 430 ms, 50.1% of views LCP-poor. Homepage 2,305 ms and /pricing 1,290 ms are green
17Core Web Vitals — desktop FIXhigh LCP 4,288 ms, CLS 0.16; TTFB 0.11–0.42 s, so the cause is client-side JavaScript, not the server
18noindex inside the sitemap FIXhigh 13 URLs (7 core/blog + 6 project); 4 are money pages, worst /solutions/software-localization
19Broken internal links FIXhigh 29 broken targets: 28 on the retired store.crowdin.com, 1 slug typo crowdinn on /customers
20Internal links to redirects FIXmedium 162 internal links land on a redirect; 43 chains of 2–4 hops, no loops
21Sitemap hygiene FIXmedium 249 URLs canonicalise away because of uppercase slugs; lastmod missing on 48 of 49 core; 2 duplicate <loc>
22Orphan pages FIXcritical All 19,345 /project/* reachable only via sitemap, plus 3 orphan money pages
International configuration
23Language subdomains indexable FIXcritical 18 locales return 200 with noindex, nofollow — 18 language versions do not exist for search or for models
24hreflang validity FIXhigh 48 core pages publish 22 alternates including invalid br, xh, zu, while pt-BR is missing
25hreflang coverage of the blog FIXhigh 349 blog URLs carry no hreflang at all — the whole editorial section sits outside the international config
26Locale sitemaps FIXmedium Every locale map returns main-domain URLs; xh/zu serve HTML with code 200 where XML is expected
27Regional meta differentiation Blocked Not assessable while all 18 locales are noindex; re-run after check 23 is resolved
AI answer readiness
28Query fan-out coverage FIXhigh 53% (49 yes / 30 partial / 41 no of 120). Self-hosted 0 of 5, comparisons 1 yes / 9 no, pricing 9 no, security 5 no
29Content extractability FIXhigh 65.1 of 100 on content present inside structures, against 88.9 for structures existing at all
30Templated thin pages FIXcritical 19,361 of 19,743 (98.1%) — 19,345 UGC plus 16 core; Jaccard median 0.993, 1 unique token
31Near duplicates PASS* 0 across all 398 core and blog pages; the 98.0% duplication is entirely the /project/* template
32Machine-readable comparison tables FIXhigh /alternatives/lokalise-alternative: 83 of 168 cells (49%) carry meaning only in an aria-hidden SVG
33E-E-A-T on money pages FIXhigh 44 of 46 below 60 (commercial median 48) against a blog median of 87 — two templates
34Author bylines PASS 311 of 311 articles (100%) have a named author, verified visible on a random sample of 40
35Author pages FIXhigh 0 pages for 42 authors; /blog/author/* returns 404. 7 pages would cover 274 of 311 articles
36Person schema completeness FIXhigh 0 of 312 nodes carry @id, url, jobTitle or worksFor; job titles sit in prose
37Author identity correctness FIXhigh Ronak Ganatra's sameAs points at Akanksha Babbar's LinkedIn — a false identity claim, not a broken link
38Freshness signals FIXmedium 0 of 10 top pages have a determinable date; 46 updates are invisible to machines; dateModified on 52 of 311
39Keyword cannibalisation FIXmedium 13 clusters / 46 pages; 7 commercial-vs-informational. translation software pulls six /solutions/* pages onto one term
40Semantic cannibalisation PASS* 5 pairs above 0.85 cosine, but literal overlap is only 0.095–0.146 — differentiate, do not redirect
41Missed internal linking FIXmedium 4 unlinked pairs in the 0.75–0.85 band (of 10 in band), plus 5 money pages with zero editorial inlinks
42Unsubstituted template variables FIXhigh {count} visible on 12 pages including the homepage, plus 3 meta descriptions and 6 JSON-LD nodes
43Placeholder copy in production FIXmedium A live H2 on /features/in-context-translations reads "Highlight the specific "Web" pain points here". Scope: 1 page of 898
44Internally consistent facts FIXmedium Language count stated as 700+ / 300+ / 100+ on different pages; a competitor price is 4 months stale
45CTR against impressions Blocked No Search Console access — a central check of this section by our methodology. Request Full or Restricted for [email protected]
46Pages-decay Blocked Same cause. Needs two non-overlapping GSC windows; not inferable from any other source
Structured data
47JSON-LD syntax validity PASS 0 parse errors across 898 pages; no competing microdata or RDFa vocabularies
48Fabricated review markup PASS aggregateRating and Review appear nowhere — Crowdin marks up no rating it does not own. Keep it that way
49FAQ markup matches visible text FIXcritical 2 pages carry a fully hidden FAQPage (7 and 4 questions, 0 visible); 4 more have single-question gaps
50Entity grounding (sameAs) FIXhigh 358 of 359 pages carry an Organization stub with no sameAs; the full node exists on the homepage only
51Graph consistency FIXcritical #organization declared twice with different fields on 311 articles — two generators in one theme
52Dangling @id references FIXhigh 43 of 48 core pages point about/mainEntity at a #webapplication node that is not on the page
53Placeholders in markup FIXhigh 6 pages ship {count}, {name} or {user_name} inside JSON-LD and the meta description
54Markup coverage of the corpus FIXhigh 359 pages have JSON-LD — 1.8% of the 19,743-URL site, because 98% is /project/* with none
Visibility in models
55Brand mentions in answers PASS* Mentioned in 11 of 12 prompts at average position 2.9 — a description of 12 answers, not a visibility rate
56Domain cited as a source FIXcritical 0 of 12 answers cited crowdin.com. Measured citations: 15 vs Smartling 203 — 9th of 10
57Presence in the category dataset FIXhigh Absent from the 42-brand localization entity set while competitors are present — verified by querying those competitors
58Positioning the model repeats FIXhigh "Open source / community / developer-heavy"; "best overall" goes to Phrase. An accurate reading of the current copy
59Bot crawl activity (proxy signals) PASS Common Crawl 1,225 apex records in July 2026 plus byte-identical live bot responses — crawling is happening
60Quantitative share of voice Blocked Requires Ahrefs Brand Radar, outside this run. We do not substitute another method — a spot-check percentage would overstate it
61LLM traffic share and ARPU Blocked No GA4 access. Six metrics unavailable: LLM traffic share, share of organic, engagement, bounce, ARPU, top landing pages
62Perplexity and Claude coverage Blocked LLM Mentions covers ChatGPT and Google AI Overview only; Perplexity spot-checks hit a Cloudflare rate limit on our IP
63Crawl activity from server logs Blocked Optional check — no logs provided. Proxy signals covered it; 30–90 days of logs would make it exact
Semantics & links
64Semantic coverage gap FIXhigh 297 keywords / 126,880 searches a month held by competitors; 89 of them belong to pages Crowdin already publishes
65Intent validity of the gap PASS A bid test removed 88,500 searches/mo (41%) of the apparent gap as non-buyer intent before any of it reached this report
66Alternatives hub coverage FIXhigh 2 pages against 13 competitors; competitor brand demand 13,510/mo at bids to $49
67Referring-domain gap FIXhigh 2,026 donors link to 2+ competitors and none to Crowdin; 1,074 editorial, 249 with rank ≥ 60
68Community and directory presence FIXhigh awesome-i18n (408★) omits Crowdin entirely while listing POEditor and naming SimpleLocalize 6 times
69Review-profile consistency FIXhigh Trustpilot 3.8 on 18 reviews against G2 4.4/662 and Capterra 4.7/175 — the spread reads as uneven quality
70Donor inventory quality PASS 890 domains excluded with a stated reason each; 24 footprint domains isolated by two signals, not spam score alone
71Reddit and community mentions Blocked 403 from our environment on all 15 subreddits — this is "not measured", not "no mentions". Needs Reddit API access
72Pitch Targets Blocked Not computable without the contractor link inventory. 645 link-gap domains supplied instead — not equivalent
73Actual impressions of the long tail Blocked No Search Console — the only source showing queries with impressions but no ranking position
73 checks: 21 PASS or PASS*, 42 FIX, 10 Blocked. Every row above is expanded in the section that follows, with the underlying rows in the client spreadsheet. A Blocked status means the check was not performed and the reason is named — it never means the result was zero. Appendix B lists exactly what to send us to close each one.

2.AI crawler access

This is the part of the audit that decides whether anything else matters, and it comes back clean. Nothing on crowdin.com is withheld from AI crawlers: they are served the same bytes a browser receives, and JavaScript hides no content from them. One subdomain throttles them, and one class of information — the price — is missing from the text layer everywhere.

URLs blocked by robots.txt
0
of 19,743 in all three sitemaps
Bot vs browser responses
identical
13 UAs × 9–10 URLs × 3 runs, byte for byte
Pages needing JavaScript
0
of 888 gain >20% text from rendering

robots.txt: nothing blocked, but six dead rules

We did not assess robots.txt by reading it. We matched every rule against every URL in all three sitemaps using Google's matching semantics, because a rule that looks alarming often applies to nothing. That is exactly what happened here. Our own opening hypothesis — that Disallow: /blog/post/* was hiding the blog — was wrong, and we disproved it before it reached this report: articles live at /blog/<slug>, and 0 of 349 blog URLs match that prefix.

SitemapURLsUnder a disallow
sitemap-core.xml490
blog/sitemap.xml3490
sitemap-community0.xml19,3450
Total19,7430 (0.00%)

There is no per-bot section anywhere on any host for GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Google-Extended or CCBot — all of them fall under User-agent: *, and * blocks no public content.

What we did find is decay. Of the 10 rules under User-agent: *, 6 point at URLs that now return 404: /app, /backend, /download, /translate, /case-studies/* and /blog/post/*. The remaining four (/join, /login, /settings, and a JS asset path) are legitimate. Today the dead rules cost nothing. They matter for two reasons: they already caused one analytical error in this audit, and at the next URL refactor a stale prefix can silently start matching live pages.

One rule is malformed rather than dead: the token User-agent: http://Mail.RU is not a valid user-agent string, so that block applies to no crawler at all. If blocking Mail.RU is intended, the token must be Mail.RU.

Live bot check: byte-identical responses

A robots.txt audit alone proves nothing — blocking usually happens at the CDN or WAF, below the file. So we requested the site as the bots themselves, comparing both status code and response size in bytes against a browser baseline. 13 AI user agents across 9–10 URLs, then a stability series of 3 runs × 7 UAs: all 63 measurements returned identical byte counts, with zero variance.

URLCodeBytesBot vs browser
crowdin.com/200893,486identical for all 13 UAs
/pricing200137,749identical, stable across 3 runs
/features/ai-translation200472,441identical
/blog/localization-strategy200134,967identical
/project/binance20015,815identical
support.crowdin.com/20041,482identical
community.crowdin.com/42966–69blocked for GPTBot and ClaudeBot
Why a green result here is evidence and not just an absence of errors. Per our own rule, "no problem found" only counts if we can say how a problem would have shown up. Blocking would have appeared as one of four signals: 403/503 on a bot UA, 503 with a cf-ray challenge header, a 200 with a much smaller body (a stub page), or a 429 rate limit. All four are observable in this method, and one of them did fire — the 429 on the forum. The check is therefore not "always green".

Caveat on CDN and WAF. We found no Cloudflare or Akamai signatures in 130+ requests (server: nginx on the apex, Vercel on support and store, no cf-ray anywhere). This is our assessment from response headers, not confirmation from Crowdin. A WAF can hold rules scoped by geography or ASN that our exit point never triggered. Please confirm what sits in front of the origin.

Common Crawl: the site is in the index, including the blog

Common Crawl is a training and retrieval corpus for many models, so presence there is a direct proxy for whether content has been collected. We queried the latest index, CC-MAIN-2026-30 (July 2026), via curl.

TargetRecordsDetail
crowdin.com/*1,2251,118 × 200 (91%); crawled 10–22 July 2026
crowdin.com/blog/*131105 successful captures — independent proof the blog is not blocked
support.crowdin.com/*426418 × 200 (98%) — documentation is collected best of all
store.crowdin.com/*324314 × 200
community.crowdin.com/*73single day only — corroborates the 429 throttle below
control: en.wikipedia.org/*72,852confirms the query itself returns large result sets
A methodological trap worth recording, because it nearly produced a false result: in this environment Python's urllib fails with CERTIFICATE_VERIFY_FAILED against the Common Crawl API. Had we accepted that, every target above would have read zero records — indistinguishable from "the site is not in the index". We switched transport to curl. A tool failure recorded as a zero is the most damaging error available to an audit, which is why the Wikipedia control row exists.

Content without JavaScript: no barrier on any platform

crowdin.com is not one system. We identified five stacks and checked each for whether content survives without JavaScript, since many AI crawlers do not execute it.

SectionPlatformReadable without JS
crowdin.com/, /features/*, /blog/*Astro 7.0.6 (static)Yes — homepage 2,256 words and 13 H2 in raw HTML
/pricingAstro 7.0.6Yes (974 words) — but no price figures
/project/*legacy Crowdin appYes — 12 of 12 in a random sample
support.crowdin.comAstro 7.1.6 / StarlightYes — 459–9,708 words (n=8)
store.crowdin.comNext.js on VercelYes, server-rendered (n=5)
community.crowdin.comDiscourseYes — identical 3,381 characters to bot and browser

A raw-versus-rendered comparison first suggested a problem: the homepage went from 2,256 words to 3,999 after rendering, /pricing from 974 to 3,799. We checked what the extra text actually was, and it is the cookie-consent banner expanding into roughly 2,000 words of cookie tables. No substantive content block exists in the rendered DOM and is missing from raw HTML. The ratio was a false signal; we report the verified conclusion.

Findings
  1. Nothing is blocked. robots.txt disallows 0 of 19,743 sitemap URLs; no host carries a rule naming any LLM bot.
  2. No user-agent discrimination. 13 AI bots receive byte-identical responses to a browser across 9–10 URLs and 3 runs; 63 of 63 stability measurements matched exactly.
  3. The site is being collected. 1,225 Common Crawl records on the apex and 131 on the blog in July 2026 alone.
  4. JavaScript is not a barrier on any of five platforms; 0 of 888 pages gain meaningful text from rendering.
  5. High — the forum throttles AI bots. community.crowdin.com returns 429 to GPTBot, ClaudeBot and Claude-SearchBot with discourse-rate-limit-error-code: gptbot_crawler_rate_limit. A browser UA made 6 consecutive requests with no pause and got 6 × 200, so this is bot-specific, not general load shedding. Observed ceiling is roughly one request per 30 seconds. Corroborated independently: 73 Common Crawl records against 426 for the documentation site. This is Discourse's built-in limiter, not a CDN or WAF.
  6. Medium — the price is not in the text layer. A regex for currency figures over the rendered /pricing body returns an empty set, and over the entire HTML returns 0. This is not content behind JavaScript; the figures are not present as text at all.
  7. Low — six dead rules in robots.txt point at 404s, plus one malformed User-agent: http://Mail.RU token that applies to nothing.
What to do

1. Stop throttling AI bots on the forum (client, one setting). In Discourse, remove the LLM user agents from slow_down_crawler_user_agents and review slow_down_crawler_rate / crawler_rate_limit_*. Target at least one request per second for GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot and CCBot. The forum is user-generated answer content — precisely the material models quote.

2. Publish the price as text (client, Sprint 2). Server-render the plan tiers so the figures exist in HTML. Detail and the necessary caveat are in section 4.

3. Delete the six dead robots.txt rules and fix the Mail.RU token (client, minutes). Hygiene, not a visibility gain — do it because a stale prefix becomes dangerous the moment URLs change.

4. Confirm what CDN or WAF fronts the origin (client, one answer). Our finding is header-based; a geo- or ASN-scoped rule would be invisible to us.

▧ Full crawl, all checked URLs · 898 rows

3.Technical audit

The technical fundamentals are in better shape than the score suggests. Response codes, canonicals, HTTPS, HSTS, title uniqueness and the 404 handler are all clean — 898 crawled URLs return 200, and not one page has a duplicate title or a second H1. The crawler's health score of 42 is driven by three things instead: page speed on the application, an indexation structure that publishes 19,345 near-identical URLs, and an international setup where 18 language versions are switched off.

Health score
42/100
Screaming Frog, 898 URLs in list mode
Mobile LCP, field data
6,815 ms
50.1% of views rated poor (threshold 2,500 ms)
Duplicate titles & H1
0
of 898 pages — unusually clean

What passes

Worth stating plainly, because these are the checks that are expensive to fix later and are already right: every one of the 898 sitemap URLs returns 200, with no 4xx or 5xx among them. /this-page-does-not-exist returns a real 404 with a branded page rather than a soft-404. http redirects to https and www to non-www, both 301, with HSTS at max-age=31536000; includeSubdomains. There are 0 canonical conflicts across all 898 pages and 0 missing canonicals on core. All 898 titles are unique, there are 0 duplicate H1s and 0 pages carrying two H1s. Near-duplicate content across the 398 core and blog pages is 0.

Critical — page speed on the application, not the marketing pages

Field data from real Chrome users (CrUX, 28 days to 15 Aug 2026) puts the origin outside threshold on three of four metrics on mobile: LCP 6,815 ms, INP 430 ms, FCP 6,718 ms, with 50.1% of views LCP-poor. Desktop is also failing: LCP 4,288 ms and CLS 0.16.

The aggregate hides where the damage is, which is why we broke it out by URL. Only five URLs have enough traffic for their own CrUX record, and the split is decisive.

URLField LCPVerdict
/ homepage2,305 msgreen — within threshold
/pricing1,290 msgreen — the fastest page measured
/settings5,986 mspoor for 48% of views
/project/*up to 8,021 mspoor — 98% of the site by URL count
/profile8,804 mspoor for 64% of views — the worst measured

So the marketing pages that matter for AI visibility are fine, and the origin-level failure is produced by the logged-in application. Two facts identify the cause precisely. Server response time across all 898 pages is 0.11–0.42 s, so the backend is fast. And lab testing of /project/ankidroid finds 1,920 KiB of unused JavaScript on a page whose entire text content is 562 words, with 7.2 s of main-thread work. The problem is client-side JavaScript weight, not infrastructure.

Per-URL field data does not exist for the core and blog pages: CrUX withholds records below a traffic threshold. That is a property of the data source, not a measurement failure and not evidence those pages are fast. Our verdict rests on CrUX field data; PageSpeed Insights lab scores are used only to identify causes, never for the verdict. Divergence between lab and field by a factor of several is normal.

Critical — 18 language versions are switched off

Crowdin publishes 21 language subdomains. We checked 20 of them on /pricing. 18 return HTTP 200 together with noindex, nofollow. At the same time 48 core pages publish 22 hreflang alternates pointing at those very URLs — the site tells search engines and models that translated versions exist, then forbids every one of them from being indexed. Localization spend is not converting into visibility.

Two of the 21 are not language versions at all. xh.crowdin.com and zu.crowdin.com return 26,346 bytes of identical HTML titled "Crowdin Enterprise" — a single-page app for a different product, with no hreflang, no canonical and no meta robots. Their /sitemap.xml returns an HTML document with status 200 instead of XML, so a crawler asking for the sitemap gets an app page and is never told anything went wrong. Separately, every locale's sitemap returns the main-domain URL set, so even if noindex were lifted tomorrow the locale sitemaps would cover nothing.

High — four money pages are excluded from the index by their own markup

13 URLs sit in the sitemap carrying noindex — 7 core and blog pages plus 6 closed projects. The site simultaneously requests indexing and forbids it. We confirmed each with a browser user agent, so this is served to everyone and is not a crawler artifact. Four are money pages.

URLDirectiveWhy it matters
/solutions/software-localizationnoindex, followTargets localization platform — 140/mo at a $39.06 top-of-page bid, the most expensive term in the niche
/features/quality-assurancenoindex, followBest-positioned page for app translation (2,900/mo) — see cannibalisation in section 4
/features/workload-managementnoindex, followFeature page, 0 editorial inlinks
/learnnoindex (no follow)Educational hub, E-E-A-T 62 — second-best commercial page on the site
▧ All 13 noindex URLs found in sitemaps · 13 rows

High — broken links and stale redirects

29 internal link targets return 404. The cause is a single retired system: 28 of 29 point at store.crowdin.com app pages linked from "What's new at Crowdin" posts published between 2020 and 2026, where the apps were renamed or removed and the links never updated. The 29th is a slug typo — crowdinn with a double n — and it sits on the money page /customers. Note that the crawler's own inlink export reported zero broken links; that was a list-mode artifact, so we extracted 3,154 link targets from raw HTML and checked the 1,314 Crowdin-domain ones directly.

Beyond outright breaks, 162 internal links land on a redirect rather than the final URL, and 43 chains run 2 to 4 hops (no loops). Common cases: /ai-localization/features/ai-translation, /alternatives/, and legacy dated blog URLs such as blog.crowdin.com/2021/08/25/translation-memory/ taking three hops.

▧ Broken links with source pages · 41 rows

High — hreflang is inconsistent, and the blog has none

TemplatePageshreflang state
core4822 codes — includes invalid br, xh, zu; pt-BR missing
/project/*50020 codes — correct pt-BR; no br/xh/zu
blog349no hreflang at all

Two things stand out. br is the code for Breton, not Brazil; given the URL br.crowdin.com the intent was Brazilian Portuguese, whose correct code is pt-BR — and the /project/* template already uses it correctly. So the marketing pages advertise a Breton edition that does not exist while never declaring the Brazilian one. That the two templates disagree is itself the diagnostic: hreflang changes are shipping out of sync, and the core template is running an older version. Meanwhile the entire blog — the section that normally earns both search and LLM traffic — sits outside the international configuration completely.

Medium — sitemap hygiene

249 URLs in sitemap-community0.xml canonicalise to a different address purely because their slugs contain capital letters (/project/My_F-droid, /project/Nyachi). No lowercase twin appears in the sitemap, so the file systematically submits the non-canonical variant. Also: lastmod is absent on 48 of 49 core URLs, and there are 2 duplicate <loc> entries.

Findings
  1. Critical — field Core Web Vitals fail at origin level: mobile LCP 6,815 ms, INP 430 ms, 50.1% of views LCP-poor; desktop LCP 4,288 ms, CLS 0.16. The homepage (2,305 ms) and /pricing (1,290 ms) are green — /profile (8,804 ms), /settings and /project/* drag the origin down. TTFB is 0.11–0.42 s, so the cause is client JavaScript: 1,920 KiB unused on a 562-word page.
  2. Critical — 18 language subdomains return 200 with noindex, nofollow while 48 core pages publish 22 hreflang alternates pointing at them. xh and zu are a different product's SPA, and their sitemap returns HTML with status 200.
  3. High — 13 sitemap URLs carry noindex, four of them money pages. The worst is /solutions/software-localization, targeting the niche's most expensive term at a $39.06 bid.
  4. High — 29 broken link targets: 28 on the retired store.crowdin.com from "What's new" posts, 1 a crowdinn typo on /customers.
  5. High — hreflang is inconsistent: 349 blog URLs have none; core publishes invalid br/xh/zu and omits pt-BR, which the /project/* template gets right.
  6. Medium — 162 internal links point at redirects and 43 chains run 2–4 hops.
  7. Medium — 249 sitemap URLs canonicalise away because of uppercase slugs; lastmod missing on 48 of 49 core URLs.
What to do

1. Lift the four money pages out of noindex (client, one commit). Highest value per minute in this section: these pages exist, are written, and are currently invisible.

2. Decide on the 18 locales (client, decision first). Either publish them — remove noindex, generate real localized sitemaps — or retire them and strip the 22 hreflang alternates from the core template. The current state is the only option with no upside. Separately, remove xh and zu from hreflang: they are a different product.

3. Split the speed problem from the marketing site (client, Sprint 2). The fix is JavaScript weight on /profile, /settings and /project/* — or moving the app to its own origin so its metrics stop being attributed to crowdin.com. Do not optimise the homepage or /pricing; they already pass.

4. Fix 29 broken links and 162 redirect links (agency, mechanical). One find-and-replace retires most of the store.crowdin.com references.

5. Add hreflang to the blog and correct the codes on core (client). Replace br with pt-BR, reusing the /project/* template that is already correct.

▧ Prioritised fix list, all sections · 77 rows

4.Readiness for AI answers

This is the core of the audit. Models can reach every page on crowdin.com; the question this section answers is what they find when they get there. The short version: the blog gives them a great deal, and the pages that sell give them very little. Two numbers frame everything below — a model asking the twelve follow-up questions a buyer asks finds answers to 53% of them, and 44 of 46 commercial pages score below 60 on E-E-A-T against a blog median of 87.

Query fan-out coverage
53%
49 yes / 30 partial / 41 no across 120 sub-questions
Content extractability
65.1
of 100 — facts present inside structures, top 10 pages
Templated thin pages
98.1%
19,361 of 19,743 URLs in the sitemap

Critical — 98% of the site is one page published 19,345 times

The /project/* catalogue is 19,345 URLs, 98.0% of the site, and it is indexable, in the sitemap, and statistically a single document. We measured sameness three independent ways and all three agree:

MeasurementResultMethod
5-gram Jaccard similaritymedian 0.993–0.995Two independent samples; 400/400 and 780/780 pairs above 0.98
Unique tokens per pagemedian 1Range 0–4 — the project name, and nothing else
Crawler near-duplicate score100 on 500 of 500Screaming Frog at a 90% threshold
Word count562 on every pageZero spread — identical interface text, not content
H1 and structured data0 and 0No heading, no markup, on all 500 sampled

Note the trap in that table: 562 words clears any sane thin-content threshold, so a word-count test would clear these pages. They are thin by substance, not by length.

They are also unreachable except by sitemap. The /projects listing emits zero links to any project page in raw HTML — the list is built client-side, and ?page=1 and ?page=2 return byte-identical HTML at 173,953 bytes. The crawler appears to show 6,490 inlinks to these pages, but every one is a self-link from the page's own hreflang and in-app navigation: cross-page links are 0, and no core or blog page links to a single project. All 19,345 are true orphans.

Why this is one decision and not 19,345 tasks. The pages are generated from one template and submitted by one sitemap file. Setting noindex, follow on that template and removing sitemap-community0.xml from the sitemap index retires the entire class in a single change — and follow preserves whatever link value the pages pass. This is not a recommendation to delete anything: the projects stay live and usable for their users; they simply stop being submitted as indexable marketing pages.

Critical — the price exists nowhere a model can read it

Crowdin's pricing is published only inside JSON-LD markup. We verified this from the primary source rather than a crawler report:

Where we lookedDollar figures found
Visible text on /pricing0
The entire /pricing HTML, scripts and attributes included0
Rendered DOM after network idle0
JSON-LD Offer nodes0 / 50 / 150 / 450 USD

The "Compare Plans" table is empty in the markup: all 42 of its cells contain no text, each holding a loading skeleton instead — the page carries 69 animate-pulse placeholders and 23 hydration islands. Crucially this is not fixable by rendering: the crawler ran with JavaScript enabled and still recorded no prices, because the table hydrates from data that never arrives inside the render window. The plan names "Pro" and "Team+" likewise appear only inside the markup; visible text has only "Free plan" and "Enterprise".

What we are not claiming. We are not saying the marked-up figures 0/50/150/450 are correct. The primary source is closed: the pricing endpoint returns {"success":false,"acl_error":true} without authentication, and the JavaScript bundle contains an OldPricingNotice component that checks whether a plan is on new pricing — direct evidence that tiers have changed at some point. The verified finding is narrower and still serious: there is no source for the price in the text layer. Confirming the figures themselves needs the current price list from Crowdin, so that item is Blocked.

The practical consequence: asked "how much does Crowdin cost", a text-reading model gets nothing from the pricing page and will either stay silent or take the number from a third-party comparison article — that is, from a competitor's page. The only dollar figures readable anywhere in the top ten are inside the /alternatives/* tables.

Critical — two pages carry FAQ markup for questions that are not on the page

Structured data must describe what a visitor sees. On two pages it does not:

URLQuestions in markupQuestions visibleSeverity
/solutions/game-localization70Critical — policy risk
/alternatives/lokalise-alternative40Critical — policy risk
/blog/ai-localization87Medium — single gap
/blog/meet-crowdin-copilot76Medium — single gap
/blog/mt-post-editing43Medium — single gap
/blog/multilingual-marketing87Medium — single gap

On the first two, the words "FAQ" and "Frequently" do not appear on the page at all — we checked the live response, 434 KB and 220 KB respectively. Corpus-wide this is 15 of 285 questions (5.3%) not visible, and 11 of those 15 fall on those two pages. The remaining four are single omissions on blog articles and are ordinary content debt.

Worth knowing before prioritising: across all 898 pages Google generates no FAQ rich result whatsoever for this site. So the hidden markup carries policy risk with zero search upside; its only remaining value is LLM extraction, which is exactly why the questions belong in the visible text.

High — a model asking a buyer's questions finds 53% of the answers

When a model is asked to recommend a localization platform it does not read one page; it decomposes the request into follow-up questions and looks for each. We ran the twelve questions a real buyer asks against each of the top ten money pages — 120 checks — and graded each by reading the page.

Question typeYesPartialNoReading
HOW — how do I actually do it1414Strongest area — mechanics are well covered
INT — integrations and stack fit821Strong
DEF — what the product is730Strong
PRF — proof and numbers423Adequate
LIM — limits and constraints373Weak — mostly vague
ICP — who it is for232Weak
PRC — pricing729Failing
SEC — security and compliance225Failing
CMP — comparison with alternatives169Failing
SELF — self-hosted and on-premise005Total absence

Read the pattern rather than the total. Crowdin covers how the product works and answers poorly on why to choose it — price, security, comparison and fit. Those are the selection criteria, which is precisely why the model names Crowdin in 11 of 12 answers and recommends it in none. SELF at 0 of 5 reproduces the one prompt where Crowdin was absent entirely: asked for a self-hosted Lokalise alternative, the model returned Tolgee, Weblate, Traduora and Pontoon. The words "self-hosted", "on-premise" and "private cloud" appear on none of the ten money pages, even though Crowdin Enterprise offers it.

The first measurement of this was wrong and we discarded it. An automated keyword pass scored coverage at 66%. Reviewing the evidence string behind every credit showed the matches were landing on interface furniture: on /, "how much does it cost" matched inside a customer testimonial reading "Super optimal pricing"; on /solutions/software-localization it matched the "Start Free Trial" button; on /features, "who is it for" matched a "Try Crowdin Enterprise instead" button. After manually reviewing all 120 evidence strings, coverage fell to 53%. We report 53%; the per-row reason is in the fan-out tab.
▧ All 120 sub-question checks with evidence · 120 rows

High — the main comparison table is invisible to text readers

/alternatives/lokalise-alternative is one of the two most important pages for AI visibility, because "X vs Y" and "X alternatives" are exactly the queries models fire when asked to recommend a tool. Its comparison tables have 168 cells, of which 83 (49%) contain no text at all — the value is carried by an inline SVG marked aria-hidden="true", with no label and no title: 55 checkmarks and 28 crosses, plus 97 aria-hidden attributes on the page. Extracted as text, a feature row reads:

| Continuous AI fine-tuning |  |  |

Two whole tables — "Similarities Between Lokalise and Crowdin" and "Crowdin Enterprise Unique Features", 15 rows each — contain not one value in text form. A model reading this page cannot tell who wins any comparison.

The fix is already written. The neighbouring /alternatives/phrase-alternative uses the same component and renders 40 of 43 cells as text — only 3 are icon-only, and it carries 18 aria-hidden attributes against 97. It spells values out in words ("Available on all plans at no extra cost from us") instead of drawing a checkmark. So this is one component change, reusing a pattern that already works on the adjacent page.

High — the money pages carry no authorship, dates or sources

Across the whole corpus, E-E-A-T scoring splits cleanly by template:

Blog articles (median, n=236)
87
Other pages (median)
84
Commercial pages (median, n=46)
48
UGC /project/* (median)
5

44 of 46 commercial pages score below 60, with a median of 48, against a blog median of 87. Only two clear 60: the homepage at 88 and /learn at 62. The signal string attached to nearly every failing page is identical — no author; external links 0; no dates — and the pages sharing it are the /features/* and /solutions/* templates. The same repeated cause means the same single fix: adding an author or reviewer, a visible date and outbound references to the two templates raises all 44.

What makes the gap notable is that Crowdin clearly knows how to do this. The blog carries a named author on 100% of 311 articles. The capability exists; it just stops at the templates that sell.

High — three money pages have no links pointing to them

Three commercial pages are orphans: exactly one file in the entire 898-page crawl links to each of them, and that file is the page itself. Not the homepage, not /features, not /solutions, not any article.

URLPriorityTarget keywordVolumeBid
/alternatives/lokalise-alternative69lokalise1,000/mo$19.64
/alternatives/phrase-alternative59phrase alternative260/mo$18.46
/features/figma-plugin59figma translation

Two of the three are the alternatives pages — the highest-value comparison content Crowdin owns, and the exact page type models look for. A further two money pages (/features/workload-management and /solutions) have inlinks from navigation only and zero editorial links.

▧ Orphans and missed internal links · 7 rows

High — a template variable is printed on the homepage

{count} appears in the visible body text of 12 pages, including the homepage, where a visitor and a model both read:

You have content in {count}+ disconnected tools (GitHub, Figma, CMS)

The full list: /, /features/quality-assurance, /features/translators-workbench, /learn/continuous-localization-ebook, /learn/mobile-app-localization-ebook, /product/agency-partners, /solutions/customer-support-translation, /solutions/e-commerce-translation, /solutions/game-localization, /solutions/software-localization, /solutions/website-translation and /webinar/mobile-app-localization. It also leaks into 3 meta descriptions and 6 JSON-LD nodes (where {name} and {user_name} also appear) — three layers of the same unsubstituted variable.

The damage is specific. Every one of these slots is where a number belongs: "100+ file formats", "40+ MT engines", "700+ integrations". A model extracting facts gets a broken token exactly where the quantified claim should be — the sales argument is lost precisely where it was meant to land.

A false positive we excluded. A naive search across 898 files flags 38 pages, because articles about internationalization legitimately contain {count}, {name} and {locale} as ICU MessageFormat and i18next syntax inside code examples — correct, on-topic content. We stripped <pre> and <code> blocks before counting and excluded pages such as /blog/react-i18n and /blog/flutter-localization-guide. The figure of 12 is text outside code samples only.

One related item, separate because it is genuinely isolated: a live H2 on the money page /features/in-context-translations reads "Highlight the specific "Web" pain points here" — a copywriter's instruction shipped to production. We checked scope across all 898 pages: this phrase appears on 1 page, and Lorem ipsum, TODO and FIXME appear on 0. A single incident, not a release process problem.

High — entity data is published on one page out of 359

Organization markup is present on 359 pages and, to Crowdin's credit, all of them use one stable identifier — crowdin.com/#organization, with no competing variants. But sameAs exists on exactly one page: the homepage. The other 358 carry a stub of name, URL and logo with no sameAs, no contactPoint, no address and no founding date. Measured against the full site, the company's verifiable identity is published on 1 URL in 19,743.

The cause is mechanical and therefore cheap: only 4 unique payloads exist across the entire site, byte-identical within each template, and the same stub appears on the locale subdomains too. The defect is in a shared Astro partial, not in individual pages — one edit closes all 358.

Two further graph defects compound it. Critical: on 311 articles #organization is declared twice with different fields — one node carrying description, the other legalName — because two markup generators in the theme both emit it; #website is duplicated the same way. A consumer consolidating that graph gets a non-deterministic result. High: on 43 of 48 core pages, about and mainEntity point at a #webapplication node that is not present on those pages — a dangling reference.

High — 42 authors, zero author pages

Authorship is Crowdin's strongest E-E-A-T asset and its most under-published. 311 of 311 articles (100%) have a named author, confirmed visible on the page in a random sample of 40. There are 42 authors. There are 0 author pages: /blog/author/* returns 404, and no such URL appears anywhere in the crawl.

The markup is equally incomplete. Across 312 Person nodes, the fill rate for @id, url, jobTitle, worksFor, image and knowsAbout is 0%. Job titles exist — as prose inside the bio ("currently leads a marketing team at Crowdin") rather than as a field a machine can read. So the author is a name, not an addressable entity.

Coverage is cheap to fix: 7 author pages would cover 274 of 311 articles (88%) — Diana Voroniak (108 articles), Yuliia Makarenko (70), Iryna Namaka (39), Khrystyna Humenna (25), Julia Herasymchuk (15), Yana Feshchuk (7), Andrii Bodnar (2).

One item needs correcting before any other authorship work, because it is not a broken link but a false statement about a person: on /blog/translation-quality-without-speaking-language the author Ronak Ganatra's sameAs points at linkedin.com/in/akanksha-babbar-14583420/ — a different person, who writes a different article on the same site. The markup asserts that two people are one.

▧ All 42 authors with article counts · 42 rows ▧ Every structured-data defect by page and field · 1,599 rows

Medium — 46 content updates are invisible to machines

Freshness is a ranking and citation factor for models, and Crowdin is losing credit for work it has actually done. 98 articles display "Last updated: <date>" to readers. Only 52 also carry dateModified in markup. For the other 46, the sitemap reports the original publication date. The clearest example:

URLDate shown to readersDate sent to machines
/blog/crowdin-for-figma-design-and-prototype-for-multiple-markets2025-06-102020-02-03
/blog/data-driven-approach-to-translation-quality-evaluation2026-06-292020-12-10

A five-year-old timestamp on freshly updated content is worse than no timestamp. Separately, 0 of the 10 top money pages have a determinable date by any source, and lastmod is missing from 48 of 49 core sitemap URLs. Two smaller markup defects sit alongside: unrendered markdown leaked into 293 Person.description fields, and 6 pages ship placeholder tokens inside JSON-LD.

Why we ignored the HTTP Last-Modified header. It is populated on 397 pages, but holds only 8 distinct values, all within a 31-second window matching the moment of our crawl — it reports CDN cache time, not content age. Had we trusted it, 397 pages would have been scored "fresh" and this entire finding would have disappeared. We also checked for the opposite failure, mass date re-stamping, and found none: sitemap dates hold 293 distinct values across 312 pages.

Two correct figures for article age, measured differently: 143 of 311 articles are over two years old by sitemap lastmod, but only 126 when the freshest available signal is used. The 17-article gap is the invisible-updates finding above.

Medium — competing pages and contradictory facts

13 keyword clusters spanning 46 pages compete internally, and 7 of those pit a commercial page against an informational one. Two matter commercially:

KeywordVolumePagesThe problem
app translation2,900/mo2The best-positioned page, /features/quality-assurance, is under noindex — only a blog article with priority 39 actually competes
translation software1,000/mo7Six /solutions/* money pages target one generic term at near-identical priority instead of their own verticals

The second is really a differentiation problem. The /solutions/* pages share a template, so the generic category term surfaces on all of them rather than "ecommerce translation software" or "game localization services".

We also found 5 page pairs above 0.85 semantic similarity — and this is a case where the obvious action is the wrong one. Literal text overlap on those pairs is only 0.095–0.146: they discuss genuinely different subjects (CakePHP versus Django localization; academic versus open-source programmes) in near-identical phrasing. Do not merge them with redirects or canonicals. They need differentiated copy. Merging would delete real content.

Finally, facts contradict each other across pages, and any of these could end up quoted:

ClaimWhereConflict
Number of languages/solutions/mobile-app-localization-services says 700+; /features/ai-translation says 100+ in the body and 300+ in its own titleThree incompatible figures; 700+ is almost certainly the integration count leaking into the languages field
Number of integrations/alternatives/lokalise-alternative: "Integrations | 2 | 2" two rows above "Apps Marketplace | 700+ | 50+"The same table contradicts itself
Competitor price/alternatives/phrase-alternative stamped "as of April 2026", quoting $150 against $1,045/moFour months stale on a page making specific competitor price claims — a reputational and legal exposure, not just an SEO issue
▧ Keyword and semantic overlap, all clusters · 40 rows

Two extractability numbers, and which one to act on

Two of our checks scored the same ten pages on the same 0–100 scale and disagreed: 88.9 and 65.1. Rather than average them or pick the friendlier one, we reread the pages. Both are correct, and they measure different things.

ScoreWhat it measuresUse it for
88.9Whether extractable structures exist — headings, lists, tables, FAQ markupPotential of the current structure
65.1Whether those structures contain facts — cell values, numbers, real answersPrioritising work

/pricing shows why the gap exists. On structural signals it scores 90: one table, six H2s, nine marked-up questions. Read it and all 42 table cells are empty, behind 69 loading skeletons, with the prices only in JSON-LD and the FAQ questions absent from the text. The container is there; the content is not. The same pattern hits /alternatives/lokalise-alternative (92 structural, 54 on content). We use 65.1 throughout this report and treat 88.9 as the ceiling the existing structure would reach once filled.

▧ Top 10 money pages, all scores · 10 rows

Blocked in this section

Two central checks of this section were not performed. By our methodology the CTR gap (pages that models and search engines show but users do not click) and pages-decay (impressions declining page by page over two windows) are core to assessing answer readiness. Both require Google Search Console, and we had no access. They are not reported as zero, and no other source substitutes for them — third-party estimates give modelled traffic, not actual impressions. To close both: add [email protected] in Search Console under Settings → Users and permissions, with Full or Restricted access. Money pages ranked by actual revenue are likewise unavailable without GA4; we ranked by advertiser bid instead, which is a proxy.
Findings
  1. Critical — 19,345 /project/* pages (98% of the site) are one document. Jaccard median 0.993–0.995, median 1 unique token, 562 identical words, 0 H1, 0 markup, and 100% orphaned. One template directive plus one sitemap removal retires the class.
  2. Critical — the price is not readable as text anywhere. Zero dollar figures in the whole /pricing HTML; 42 comparison cells empty behind 69 loading skeletons; the figures live only in JSON-LD, and their accuracy is unconfirmed because the source endpoint is behind authentication.
  3. Critical — FAQ markup describes questions that do not exist on /solutions/game-localization (7 of 7 absent) and /alternatives/lokalise-alternative (4 of 4), plus 4 single-question gaps on blog articles. Policy risk with no search upside — Google generates no FAQ rich result here at all.
  4. Critical — #organization is declared twice with conflicting fields on 311 articles, from two competing generators in one theme.
  5. High — fan-out coverage is 53%. Mechanics are covered (HOW 14 yes, INT 8, DEF 7); selection criteria are not (comparisons 1 yes / 9 no, pricing 9 no, security 5 no, self-hosted 0 of 5).
  6. High — 49% of the main comparison table is unreadable: 83 of 168 cells carry meaning only in an aria-hidden icon. The neighbouring page renders the same component as text — one component fix.
  7. High — 44 of 46 money pages score below 60 on E-E-A-T (median 48) against a blog median of 87, all from the same two templates lacking author, dates and sources.
  8. High — 3 money pages are orphans, including the Lokalise comparison page (priority 69), plus 2 more with navigation-only links.
  9. High — {count} is printed on 12 pages including the homepage, plus 3 meta descriptions and 6 JSON-LD nodes, in exactly the slots where a number belongs.
  10. High — entity data is published on 1 page of 359. 358 carry an Organization stub with no sameAs; the defect is one shared Astro partial.
  11. High — 42 authors, 0 author pages; 0 of 312 Person nodes carry @id, url, jobTitle or worksFor; one author's profile link names the wrong person.
  12. Medium — 46 content updates are invisible to machines; 0 of 10 top pages have a determinable date; dateModified on 52 of 311 articles.
  13. Medium — 13 keyword clusters across 46 pages compete internally; 5 semantically close pairs need differentiation, not redirects.
  14. Medium — the site contradicts itself on language count (700+ / 300+ / 100+) and integration count, and quotes a four-month-stale competitor price.
What to do

Do these first, because each is one change covering thousands of URLs.

1. noindex, follow on the /project/* template and remove sitemap-community0.xml (client, one deploy). Retires 19,345 thin orphans. The projects stay live for their users.

2. Add sameAs to the shared Astro Organization partial (client, 7 lines). Wikipedia, Wikidata, GitHub, G2, Capterra, Trustpilot and Product Hunt — 358 pages grounded in one edit. Cheapest high-impact change in this report.

3. Remove the second #organization generator (client, one edit). Fixes 311 articles.

4. Substitute {count} everywhere (client). 12 pages, 3 meta descriptions, 6 JSON-LD nodes. Start with the homepage.

5. Put text in the comparison table cells (client). Reuse the component already working on /alternatives/phrase-alternative.

6. Either delete the hidden FAQPage markup from the two pages, or publish the questions as visible text (client). Publishing is better — those are fan-out answers. Do not leave markup describing invisible content.

7. Correct Ronak Ganatra's sameAs and remove the placeholder H2 from /features/in-context-translations (client, minutes).

8. Then the template work (Sprint 2): author, reviewer, dates and outbound sources on the /features/* and /solutions/* templates (44 pages); 7 author pages covering 274 articles; dateModified plus surfacing "Last updated" into the sitemap (46 articles); server-render the price as text; reconcile the language and integration counts and refresh the competitor pricing table.

▧ Every finding with severity and action · 614 rows

5.Visibility in model answers

Crowdin is present in AI answers and almost never the recommendation. That combination — high mention rate, zero citations, never ranked first — is the most useful diagnostic in this report, because it tells us the problem is not awareness but the material available to quote.

Domain citations
15
vs Smartling 203 — 9th of 10 measured
AI search volume
643
higher than Lokalise's 441 — demand exists
Answers citing crowdin.com
0
of 12 prompts, though 11 mentioned the brand

Citation counts: ninth of ten

How often each vendor's domain is actually cited as a source in ChatGPT answers, measured on the LLM Mentions dataset for the US in English:

smartling.com
203
smartcat.com
134
weglot.com
117
phrase.com
107
centus.com
82
lokalise.com
58
localizejs.com
49
transifex.com
24
crowdin.com
15
poeditor.com
13

Set that against brand demand and the picture sharpens. Crowdin's AI search volume is 643 — higher than Lokalise's 441 and Transifex's 259. Lokalise earns 58 citations on 441 of demand; Crowdin earns 15 on 643. People are asking about Crowdin more than about Lokalise, and the models are citing it four times less. Demand is not the constraint.

A second, independent signal points the same way. Asked for the brands associated with "translation management", the dataset returns 42 brands — Phrase, DeepL, Trados, memoQ, Wordfast, OmegaT, Transifex, Localizely, Smartcat, Lokalise, TransPerfect and others. Crowdin is absent entirely. We treat that as a fact rather than a glitch for a specific reason: the query returned 42 results and Crowdin's own competitors are among them, so the dataset is neither empty nor broken. Had it failed, the competitors would be missing too.

Twelve prompts: mentioned eleven times, recommended never

We ran twelve prompts through ChatGPT in clean sessions, fixed in advance and shuffled in order, covering three types: keyword-mapped, community-style and buyer-scenario. All results are reported, including the misses.

PromptCrowdinPositionVendors in order of mention
best localization management platformYes4Lokalise, Phrase, Smartling, Crowdin, Locize, i18next
crowdin vs lokaliseYes1Crowdin, Lokalise — the only first place, on its own brand
software localization tools for developersYes3Lokalise, Phrase, Crowdin, Transifex
best translation management system 2026Yes4Phrase, Smartling, Lokalise, Crowdin, memoQ, Trados, Smartcat
continuous localization platform with github integrationYes2Lokalise, Crowdin, Phrase, Weblate
what do you use to manage translations in your appYes2i18next, Crowdin, Lokalise, Phrase
is there a good self-hosted alternative to lokaliseNoLokalise, Tolgee, Weblate, Traduora, Pontoon
how do you handle i18n strings with a remote translator teamYes3Phrase, Lokalise, Crowdin
how do we localize our React app with a team of translatorsYes3i18next, Lokalise, Crowdin, Phrase
TMS for a game studio with 12 languagesYes3Phrase, Lokalise, Crowdin, memoQ, Smartling
mobile app startup adding 5 languages, what workflowYes4Localize, Phrase, Lokalise, Crowdin
platform supporting Figma and CI/CD for a SaaS productYes3Phrase, Lokalise, Crowdin, Locize, i18next

Crowdin appears in 11 of 12 at an average position of 2.9, and takes first place only in the branded "crowdin vs lokalise". For comparison, Lokalise appears in 12 of 12 and Phrase in 10. And in none of the twelve answers was crowdin.com cited as a source — the mentions come from model memory, with no link back to the site. Crowdin is known and not quoted.

The one complete absence is the most instructive. Asked for a self-hosted Lokalise alternative, the model named Tolgee, Weblate, Traduora and Pontoon — Crowdin Enterprise offers self-hosting, and the model does not know it, because as section 4 showed the phrases "self-hosted", "on-premise" and "private cloud" appear on none of the ten money pages.

The positioning the model repeats is the positioning the site publishes

Across the twelve answers the model consistently frames Crowdin as "open source, community, developer-heavy". Verbatim from one answer: "Choose Crowdin if your content lives in GitHub, GitLab, or similar development workflows, or if you rely on community translation", with the table row reading "Software + open-source/community". "Best overall" goes to Phrase; "best developer-first" to Lokalise. In the SaaS scenario Crowdin gets "Good for larger translation programs" while Phrase gets "Excellent" and the "best overall" verdict — a difference of adjectives that decides the shortlist.

This is not a distortion to be corrected with PR. It is an accurate paraphrase of the pages. /solutions/software-localization carries the H2 "Built by developers, for developers"; /pricing has a dedicated "Crowdin for Open Source" section; /product/for-open-source and /product/for-academic exist as standalone pages. The phrases "for product teams" and "for SaaS companies" appear on none of the ten money pages. The model is reading the site correctly and reporting what it says. That means the label is fixable only by changing the copy — and, encouragingly, that the copy is the whole of what needs changing.

One bright spot to build on: where Crowdin is cited, the blog does the work. Of the 15 prompts in which a Crowdin URL was cited at all, four cite a real page — three of them blog articles ("is crowdin a cat tool?" cites /blog/cat-tools; "what are the three types of localization?" cites /blog/software-localization). The fourth is /project/minecraft, a lucky win on someone else's brand out of 19,345 UGC pages. The editorial content earns citations; the commercial pages do not.

Four limits on this section, all of which change how the numbers should be read.
All six GA4 metrics are Blocked. Actual LLM traffic share, share of organic, engagement rate, bounce rate, ARPU and top LLM landing pages all require the client's GA4 property, and we have no access. We can confirm the traffic physically exists — ChatGPT appends ?utm_source=chatgpt.com to the URLs it cites, so those clicks land and are separable in GA4 as a chatgpt.com referral. The volume is not measured, and is not reported as zero. Grant Viewer access to [email protected] and send the numeric Property ID; that single action clears all six.
Findings
  1. Crowdin is cited 15 times against Smartling's 203 — ninth of ten vendors measured, below Lokalise's 58.
  2. Brand demand is higher than Lokalise's (AI search volume 643 vs 441) while citations are four times lower. The gap is supply of quotable material, not awareness.
  3. Absent from the 42-brand category dataset for "translation management" while competitors are present — verified as a fact, not a query failure, because the competitors returned.
  4. Mentioned in 11 of 12 prompts at average position 2.9, first in only the branded one, and cited as a source in none.
  5. Invisible for self-hosted scenarios — Tolgee, Weblate, Traduora and Pontoon win a query Crowdin Enterprise can actually serve.
  6. The model's "open source / community / developer" framing is an accurate reading of the site; "best overall" goes to Phrase. "For product teams" and "for SaaS companies" appear on none of the ten money pages.
  7. Where citations do happen, they are blog posts — 3 of 4 cited URLs. The editorial content already works.
What to do

1. Publish self-hosted and on-premise content (client, Sprint 3). There is almost no search volume for it — this is a purely AI-visibility play, aimed at a query Crowdin currently loses to four open-source projects while offering the capability.

2. Reposition the money pages (client, Sprint 3). The model repeats what the pages say. Add explicit product-team and SaaS framing alongside the open-source story rather than replacing it; the open-source association is an asset, but it is currently the only one.

3. Fill the comparison and pricing gaps from section 4 (client, Sprint 1–2). These are the fan-out categories that decide recommendations, and they are the same fixes already listed — table cells, prices as text, security and limits content.

4. Send GA4 and Brand Radar access (client, one action each). Without them, channel performance and true share of voice stay unmeasured; we will not estimate either.

6.Semantics & content plan

The single most important thing to understand about this category is that it is low-volume and high-bid. Planning content by search volume in localization software leads systematically to the wrong pages, and we can show exactly how far wrong.

Keyword gap
297
keywords / 126,880 searches per month, after intent filtering
Already have the page
89
keywords / 25,720 sit on URLs Crowdin already publishes
Removed as wrong intent
88,500
searches per month — 41% of the apparent gap

The category pattern: 70 searches a month at a $25 bid

Advertisers do not pay for clicks that do not convert. That makes the top-of-page bid an independent check on intent, and one that volume alone cannot give. The contrast is stark:

KeywordVolume/moTop-of-page bidWhat it actually is
localization management platform70$6.38–25.51The exact category term — real B2B demand
localization platform140$11.29–39.06Real B2B demand
phrase tms110up to $738.59Highest-value term found in the niche
smartling2,400$3.41–49.09Competitor brand — commercially valuable
multilingual user interface33,100$0.25–1.10Windows MUI language packs — not a buyer
glossaries27,100$0.13–0.92Dictionary lookups
ai powered translation service14,800$0.00No advertiser demand at all

Applying that as a rule — volume above 5,000 with a top-of-page bid under $2 is presumed not to be a buyer — removed 88,500 searches a month, 41% of what first looked like the gap. Our own initial analysis had multilingual user interface as the single biggest unclaimed opportunity on the site. It is Windows language packs. We withdrew it, and the control confirms the withdrawal: seeding the keyword tool with that term returns exactly two keywords, itself and "windows multilingual user interface" — it has no semantic neighbourhood in this category because it does not belong to it.

The same discipline killed a second false opportunity. types of computer-assisted translation showed 90,500 searches a month. Its bid is $0.02–0.07, and /blog/cat-tools already ranks eighth for it. Had we skipped the subtraction step, this report would have recommended building a page Crowdin already has, for a query with no commercial value, as priority one.

The honest consequence of all this filtering: our gap figure came down from 361 keywords / 234,320 searches to 297 keywords / 126,880 — roughly halved. The smaller number is the one worth acting on.

▧ Full keyword gap pool with bids and competitor positions · 411 rows

Priority 1 — 89 keywords on pages that already exist

This is the cheapest work in the report and it is not content creation. 89 keywords worth 25,720 searches a month belong to articles Crowdin has already published, and competitors outrank them. The pattern is consistent: Crowdin has the article, someone else has the position.

Existing pageKeywords it should ownWho outranks it
/blog/subtitle-translationsubtitle translator (1,600), subtitle meaning (1,000), example subtitle (720) — 21 keywords totalsmartcat p7, lokalise p6–18
/blog/wpml-crowdin-integration-for-wordpress-localizationtranslate wordpress plugin (390, bid $14.92), wordpress translation plugin (390), wordpress multilingual plugin (210) — 12 keywordslokalise p1–5, weglot p7–13
/blog/android-app-localization-tutorialdate format android (720) and 12 keywords totalphrase p13, lokalise p20
/blog/react-i18nreact-intl (480), react localization (170), react-native-localize (110) — 9 keywordslokalise p6–10, phrase p4–8
/blog/flutter-localization-guideflutter localization (320), flutter translate (210), flutter translation (210)phrase p3–9, lokalise p4–13
/blog/angular-localization-and-i18nangular languages (210), angular localization (110)phrase p4–10, lokalise p6–16

Note this survived our own filtering: it was 112 keywords in the first pass and is 58% smaller now — and it is still priority one, because strengthening a published URL costs a fraction of launching a new one.

▧ All 89 keywords mapped to existing pages · 89 rows

Priority 2 — an alternatives hub: 2 pages against 13 competitors

Crowdin publishes two comparison pages, for Lokalise and Phrase. There are 13 direct competitors. Missing: Smartling, Transifex, Smartcat, POEditor, Weglot, Centus, OneSky, memoQ, Trados, XTM and Memsource.

The case for building them is not volume — phrase alternatives is 260 searches a month. It rests on two things. First, competitor brand queries where Crowdin has no presence total 13,510 searches a month at bids reaching $49 (weglot 3,600, smartling 2,400 at $49.09, lokalise 1,000 at $19.64, transifex 390 at $22.16). Second, and more important here: "X alternatives" and "X vs Y" are precisely the fan-out queries a model issues when asked to recommend a TMS. This is the page type that gets read by machines at decision time. The two pages that exist are currently orphans (section 4), which should be fixed before more are built.

Priority 3 — the framework cluster

18 keywords validated by two or more competitors, difficulty 2–20, 3,630 searches a month. This has the highest competitor validation of anything in the set at the lowest difficulty, which makes it the most reliable demand signal we have: react-intl (480, KD 3), flutter localization (320, KD 4), flutter translate (210), angular languages (210), translate wordpress plugin (390, bid $14.92). It overlaps heavily with priority 1, which is a point in its favour — the same work serves both.

Two smaller terms deserve explicit mention because volume-based planning would discard them: lqa at 1,300 a month, KD 5, bid $22.16 and mtpe at 260 a month, KD 0, bid $8–12. Both are validated by Lokalise and Smartling, both are unmistakably ICP vocabulary, and both are ideally shaped for LLM citation — a defined process term with a precise answer. They are top-three by commercial value despite tiny volume.

#ClusterKeywordsVolume/moMedian KDMax bidLLM intent
1VS: head-to-head comparisons2219,7908$15.55yes
2EDU: concepts & guides1116,01018$14.19yes
3NAV: competitor brand queries1213,51028$738.59no
4VERT: video & subtitles1912,78016$14.76no
5DEV: framework & integration how-to6211,09016$27.80yes
6DEV: locale & encoding3111,08026$93.68no
7TMS: category & CAT tooling198,5500$738.59no
8BEST: listicles & reviews154,45035$23.95yes
9QA: quality & process (lqa)21,4705$22.16no
10AI-MT: machine translation (mtpe)71,62023$12.11no
▧ All 17 clusters · 17 rows ▧ 15-page content plan with targets · 15 rows
Three limits on this section.
Findings
  1. The category is low-volume and high-bid. localization management platform is 70 searches a month; localization platform is 140 at bids to $39.06; phrase tms is 110 at up to $738. Volume-led planning misreads this market.
  2. A bid test removed 88,500 searches a month — 41% of the apparent gap — before it reached this report, including our own former top opportunity.
  3. The real gap is 297 keywords / 126,880 searches, halved from our first pass. 62 are validated by two or more competitors.
  4. Priority 1: 89 keywords / 25,720 searches sit on pages Crowdin already publishes — React, Flutter, Angular, WordPress and subtitles. This is optimisation, not creation.
  5. Priority 2: 2 alternatives pages against 13 competitors. Competitor brand demand with no Crowdin presence is 13,510 a month at bids to $49, and this is the page type models read when asked to recommend.
  6. Priority 3: the framework cluster — 18 keywords, 2+ competitors each, difficulty 2–20. Plus lqa (1,300, $22.16) and mtpe (260, $12) as high-value small-volume ICP terms.
What to do

1. Strengthen the 89 keywords on existing pages first (agency + client, Sprint 2). No new URLs. Highest return per hour in this section.

2. Build the alternatives hub — 11 pages (agency, Sprint 3). Justified by fan-out behaviour and bid value, not volume. Fix the two existing orphans first.

3. Ship the framework cluster and the lqa / mtpe process pages (agency, Sprint 3). Lowest difficulty, strongest competitor validation, ideal citation shape.

4. Do not plan by volume in this category, and do not chase the "best X" listicles for their volumebest localization software is 20 searches a month. Build those for AI visibility only.

5. Send Search Console access. It is the only way to see the long tail of queries already earning impressions.

▧ Content plan, 15 pages prioritised · 15 rows

7.External placements

In this category models mostly cite intermediaries — Wikipedia, Reddit, roundups, developer media — rather than vendor sites. That makes external placement unusually important here, and it is where the single best effort-to-effect action in the whole report sits: one pull request.

Donors linking to rivals
2,026
link to 2+ competitors and none to Crowdin
Workable editorial targets
689
after removing 890 spam and non-editorial domains
Trustpilot rating
3.8
on just 18 reviews, against G2's 4.4 on 662

The referring-domain gap

We took eight direct competitors and pulled every domain linking to two or more of them but not to Crowdin. The pool was exhausted — all 2,247 available rows — giving 2,026 unique donors, of which 4 link to all eight competitors and 26 to six. After filtering, 1,074 are editorial, 689 have authority rank 40 or above, and 249 rank 60 or above.

Only a minority can be bought. Probing the 689 editorial donors for guest-post and contributor pages found buyability signals on 44 of them — 6.4% (a random sample of 80 predicted around 7.5%, consistent with the full sweep). The conclusion matters for budgeting: this niche monetises through editorial reviews and analyst coverage, so most of the gap closes by pitching, not purchasing.

DomainRankCompetitorsWhy it is worth a pitch
nimdzi.com466The localization industry's analyst firm — highest topical weight of any target here; participate in their research
techcrunch.com845Product and funding news — needs a PR hook
sitepoint.com734Developer media that models cite — technical i18n article
tutsplus.com834Developer tutorials — technical guide
translatepress.com776WordPress plugin integration and co-marketing
pitchbook.com757Vendor profile application
entrepreneur.com805Authored column on localization
▧ 645 outreach targets · 645 rows ▧ 44 donors with a paid placement route · 44 rows

The one-pull-request win

github.com/oh-jon-paul/awesome-i18n — 408 stars, last updated 24 July 2026 — does not mention Crowdin at all. It is the largest and only actively maintained awesome-list in the category (the runner-up has 192 stars). It already has an open section, "Apps and extensions for translation management", which lists POEditor, and SimpleLocalize appears 6 times across the file. Crowdin appears zero times. The entry format is a single line, and 4 of the last 5 pull requests were merged.

Awesome-lists are heavily represented in model training and retrieval data, which is why a one-line addition to a 408-star list is the best effort-to-effect ratio in this report. We surveyed all five relevant lists so the picture is complete:

ListStars / freshnessCrowdinAction
oh-jon-paul/awesome-i18n408★, Jul 2026absentP1 — submit a PR
mbiesiad/awesome-translations192★, Jan 2026presentKeep the entry current
maidis/awesome-machine-translation203★, 2024absentSkip — MT research, wrong fit
mrhota/awesome-i18n19★, 2024absentP3 — no TMS section exists yet
aesmi/awesome-i18n1★, stale 2021presentSkip
▧ 15 platforms assessed · 15 rows

Review profiles: Trustpilot is the liability

Crowdin's review presence is strong in two places and weak in the one that is structurally built to answer brand queries:

PlatformReviewsRatingAssessment
g2.com6624.4Strong. The lever is review volume and category placement, not a link
capterra.com1774.7Highest rating, 3.7× fewer reviews than G2 — grow it
trustpilot.com183.8The problem — a small, low-rated sample
producthunt.com15.0Dormant since 2016

Two reasons this matters more than the numbers suggest. Trustpilot is structured precisely for "<brand> reviews" queries — the URL is /review/crowdin.com and the page title is "Crowdin Reviews" — so it surfaces on exactly the query a buyer runs before deciding, whereas G2 competes against hundreds of its own comparison pages. And for a model, a small low-rated sample is more dangerous than a large one: 3.8 gets weighted alongside 4.4, and the 3.8–4.7 spread across platforms reads as uneven quality rather than as a sampling artifact.

▧ All review and directory profiles · 17 rows

What not to buy

890 domains are excluded, each with a stated reason: 578 for spam score at or above 30, 540 for authority below 20, 111 on risky TLDs, 106 on free subdomain hosts, 44 with spam terms in the hostname, and 34 as non-editorial noise (job boards, URL shorteners, CDNs).

Spam score alone was not enough, and here is the case that proves it. 24 domains link to five or more competitors at once while ranking below 40 — the signature of a link network or a scraper rather than editorial interest. Several would have passed a spam-score filter comfortably: stackreaction.com (rank 34, all 8 competitors, spam 18), changelogwp.com (rank 35, 7 competitors, spam 9), hubbry.com (rank 34, 7 competitors, spam 24). Linking to every competitor in the category is what exposes them, so we required two signals rather than one.

A second exclusion group is the opposite trap — very high authority, useless as targets: 14 Glassdoor locales at rank 85–90, indeed.com (91), tinyurl.com (90), aliyuncs.com (86, a CDN linking to all 8), and site builders like Wix and Squarespace at 89–91. Sorting purely by authority puts these at the top of the list.

▧ 890 excluded domains with reasons · 890 rows

A finding we rejected

"Crowdin is missing from Wikipedia's Translation management system article" was a finding until we read the article. A scripted check returned "crowdin: not found", which was easy to write up as a competitive gap. Reading it in full shows no vendor is named anywhere in it — there is no product list to be missing from. Adding one would draw a spam revert under Wikipedia's own policies. So this is not a gap, and pursuing it would have cost credibility for nothing. We include it because it is a good illustration of the derived signal lying and the primary source correcting it — the same discipline that removed 41% of the keyword gap and 13 points of the extractability score.

Wikipedia status for real: Crowdin has its own article and is present in Comparison of computer-assisted translation tools. Both should be kept factually current via talk pages, never self-edited.
Two items Blocked.
Reddit — not measured, and emphatically not "no mentions". Reddit's API returned 403 to every one of the 15 subreddits we checked, from both reddit.com and old.reddit.com, and our search tooling does not index Reddit. This is an environment restriction on our side. Recording it as zero would be a serious error, because Reddit is one of the most-cited sources in AI answers — a false "nobody discusses Crowdin" would misdirect the whole placement priority. Needs Reddit API access to close.

Pitch Targets — not computable as defined. The group means "donors linking to competitors that are absent from your contractors' inventory", and no contractor inventory was provided, so there was nothing to subtract. We supplied 645 link-gap domains instead, which is not equivalent: some of those almost certainly sit in a contractor's price list, where buying would be cheaper than pitching. One file from the agency closes this, with no new API spend.
Findings
  1. 2,026 donors link to 2+ competitors and none to Crowdin — 1,074 editorial, 689 at rank 40+, 249 at rank 60+, and 4 linking to all eight competitors.
  2. awesome-i18n (408★, updated last month) omits Crowdin entirely while listing POEditor and naming SimpleLocalize 6 times. One line, and 4 of the last 5 PRs were merged. Best effort-to-effect action in this report.
  3. Trustpilot shows 3.8 on 18 reviews against G2's 4.4 on 662 and Capterra's 4.7 on 177. Trustpilot is built for brand-review queries, and a small low-rated sample carries disproportionate weight with models.
  4. Only 6.4% of quality donors are purchasable (44 of 689) — this category is won by pitching and analyst coverage, not by buying.
  5. nimdzi.com is the highest-value target: the localization industry's own analyst, already linking to 6 competitors.
  6. 24 footprint domains were caught by a second signal after passing the spam-score filter — they link to 5+ competitors while ranking under 40.
What to do

1. Submit the awesome-i18n pull request this week (agency, one hour). Single line in an existing section of a 408-star list.

2. Run a Trustpilot review drive (client, Sprint 3). Moving 18 reviews to a few hundred changes both the rating and its weight. Ask satisfied enterprise customers directly.
Do not transfer G2 or Capterra scores into your own aggregateRating markup — that is a policy violation. Section 4 covers the correct approach.

3. Pitch nimdzi.com, sitepoint.com and techcrunch.com (agency, Sprint 3). Nimdzi first: analyst coverage in this category carries more citation weight than any single link.

4. Buy selectively (agency). Start with 10 domains at rank 50+ linking to 3+ competitors, manually vetted before payment. Never buy from the 890-domain exclusion list or the 24 footprint domains.

5. Send the contractor inventory (agency, one file). Converts 645 unclassified domains into real Pitch Targets at no API cost.

▧ Exclusion list — check before any purchase · 890 rows

8.Roadmap

Ordered by effect against effort, not by section. Sprint 1 is deliberately front-loaded with single changes that each cover hundreds or thousands of URLs — a template edit, a directive, one pull request. Every item names the systems it touches, the expected effect, and who does it.

Sprint 1 — quick wins
One change each, mostly one deploy. Do these first.
P1
Set noindex, follow on the /project/* template and remove sitemap-community0.xml from the sitemap index
19,345 URLs — 98% of the site — in one decision · Retires the thin-orphan class without deleting anything users rely on; follow preserves link value · Owner: client
P1
Remove noindex from the four money pages
/solutions/software-localization, /features/quality-assurance, /features/workload-management, /learn · Written pages become eligible for indexing and citation; the first targets the niche's most expensive term at a $39.06 bid · Owner: client
P1
Add sameAs to the shared Astro Organization partial
358 pages in one edit · 7 lines: Wikipedia, Wikidata, GitHub, G2, Capterra, Trustpilot, Product Hunt · Grounds the company as a verifiable entity models can resolve · Owner: client
P1
Remove the second #organization generator from the blog template
311 articles · Eliminates a conflicting duplicate declaration that makes graph consolidation non-deterministic · Owner: client
P1
Submit one pull request to github.com/oh-jon-paul/awesome-i18n
One line in an existing section of a 408-star list, 4 of 5 recent PRs merged · Best effort-to-effect ratio in this report · Owner: agency
P1
Substitute the {count} variable
12 pages including the homepage, plus 3 meta descriptions and 6 JSON-LD nodes · Restores the quantified claims exactly where a model looks for them · Owner: client
P1
Render text into the comparison table cells
/alternatives/lokalise-alternative — 83 of 168 cells · Reuse the component already emitting text on /alternatives/phrase-alternative · Makes the key comparison page machine-readable · Owner: client
P1
Remove the hidden FAQPage markup, or publish the questions as visible text
/solutions/game-localization (7 questions), /alternatives/lokalise-alternative (4) · Publishing is preferable — these are fan-out answers · Clears a policy risk that currently earns no rich result · Owner: client
P1
Fix Ronak Ganatra's Person.sameAs
/blog/translation-quality-without-speaking-language · Currently asserts a different person's LinkedIn — a false identity claim, not a broken link · Owner: client
P1
Remove the placeholder H2 from /features/in-context-translations
1 page · A live heading reads "Highlight the specific "Web" pain points here" · Owner: client
P1
Change the Discourse crawler setting on community.crowdin.com
Remove LLM user agents from slow_down_crawler_user_agents; review slow_down_crawler_rate · Stops 429 responses to GPTBot and ClaudeBot; target 1 req/sec · Owner: client
P1
Delete the six dead robots.txt rules and fix the malformed Mail.RU token
Hygiene now, a real risk at the next URL refactor · Owner: client
Sprint 2 — templates and structural fixes
Development work; each item covers a page class.
P2
Server-render the pricing block so figures exist as text
/pricing — 42 empty cells behind 69 loading skeletons · Lets models answer "how much does Crowdin cost" from your page instead of a competitor's comparison article · Requires the current price list first — the marked-up 0/50/150/450 is unverified · Owner: client
P2
Decide the fate of the 18 noindex locale subdomains
Either publish them (remove noindex, generate real localized sitemaps) or retire them and strip the 22 hreflang alternates from the core template · Also drop xh/zu, which serve a different product · The present state is the only option with no upside · Owner: client, decision required
P2
Add author, reviewer, dates and outbound sources to the money-page templates
/features/* and /solutions/*44 pages from two template edits · Raises commercial E-E-A-T from a median of 48 toward the blog's 87 · Owner: client + agency
P2
Create author pages
7 pages cover 274 of 311 articles (88%) · Turns 42 named authors into resolvable entities · Then populate Person @id, url, jobTitle, worksFor · Owner: client
P2
Emit dateModified and surface "Last updated" into the sitemap
46 articles whose updates are currently invisible to machines (text says 2025-06-10, sitemap says 2020-02-03) · Recovers credit for work already done · Owner: client
P2
Reduce client JavaScript on /profile, /settings and /project/*, or move the app to its own origin
1,920 KiB of unused JS on a 562-word page; field LCP 8,804 ms on /profile · Fixes the origin-level CWV failure · Do not optimise the homepage or /pricing — both already pass · Owner: client
P2
Add hreflang to the blog and correct the codes on core
349 blog URLs have none; replace invalid br with pt-BR on 48 core pages, reusing the correct /project/* template · Owner: client
P2
Repair 29 broken links and 162 links pointing at redirects
28 of the 29 are retired store.crowdin.com URLs — one find-and-replace · Plus the crowdinn typo on /customers · Owner: agency
P2
Strengthen the 89 keywords already sitting on published pages
React, Flutter, Angular, WordPress, subtitles — 25,720 searches a month on URLs that exist · Optimisation, not new content · Owner: agency
P2
Add internal links to the three orphan money pages and the two with navigation-only links
/alternatives/lokalise-alternative (priority 69), /alternatives/phrase-alternative, /features/figma-plugin · Plus the 4 missed adjacency pairs · Owner: agency
P2
Differentiate titles and H1s across the six /solutions/* pages competing for translation software
Target the verticals ("ecommerce translation software", "game localization services") instead of the shared generic term · Owner: agency
P2
Reconcile the contradictory facts and refresh the competitor pricing table
Language count stated as 700+ / 300+ / 100+; "Integrations | 2 | 2" beside "Apps Marketplace | 700+"; the Phrase table is stamped April 2026 · Removes a reputational and legal exposure · Owner: client
Sprint 3 — content and authority
Strategic work; runs in parallel once Sprints 1–2 land.
P3
Build the alternatives hub — 11 pages
Smartling, Transifex, Smartcat, POEditor, Weglot, Centus, OneSky, memoQ, Trados, XTM, Memsource · Justified by fan-out behaviour and bids to $49, not by volume · Owner: agency
P3
Publish self-hosted, on-premise and private-cloud content
Near-zero search volume — a pure AI-visibility play · Currently loses this query to Tolgee, Weblate, Traduora and Pontoon despite Crowdin Enterprise offering it · Owner: agency + client
P3
Reposition the money pages beyond "open source"
Add explicit product-team and SaaS framing on the 10 money pages · The model's "open source / community / developer" label is an accurate reading of current copy; changing the label means changing the copy · Owner: agency
P3
Ship the framework cluster plus lqa and mtpe pages
18 keywords at difficulty 2–20 with 2+ competitors each; lqa 1,300/mo at $22.16, mtpe 260/mo · Lowest difficulty and strongest validation in the set · Owner: agency
P3
Write introductory copy for the 10 category listing pages
Thin templates: /webinar/* 6 of 6 (median 196 words), /product/* 4 of 4, /enterprise-demo at 61 words · Owner: agency
P3
Produce case studies for the four uncovered ICP segments
Education (zero cases despite /product/for-academic), open source (newest is 2022 — the very segment models already associate with Crowdin), gaming (newest 2023), LSP (one case) · Owner: client + agency
P3
Build a case-studies hub to replace the 404
/case-studies currently returns 404 while 39 stories sit scattered across the blog under a generic template · Owner: client
P3
Run a Trustpilot review drive, then Capterra
18 reviews at 3.8 against G2's 662 at 4.4 · Trustpilot is structured for brand-review queries, and a small low-rated sample carries outsized weight with models · Owner: client
P3
Pitch nimdzi.com, then sitepoint.com and techcrunch.com
Nimdzi is the category's own analyst, already linking to 6 competitors · 689 quality donors link to rivals and not to Crowdin; only 6.4% are purchasable, so pitching is the main route · Owner: agency
P3
Pursue SOC 2 Type II certification
The only recommendation here requiring a project rather than a publication. ISO 27001 is held but unmarked; SOC 2 appears in enterprise procurement checklists and ISO 27001 does not substitute · Security scored 2 yes / 5 no in fan-out · Owner: client
P3
Mark up the unused assets
6 webinars without Event, 4 ebooks without DigitalDocument, 33 podcast episodes as plain Article, 8 named executives on /about with no Person markup · Owner: client
▧ All 39 case studies with ICP segment and quality flags · 39 rows ▧ Full prioritised fix list behind this roadmap · 77 rows
Four things not to do. Each is a plausible next step that would make matters worse.

A.Methodology & limitations

A. Methodology, sources and limitations

Crawl. Screaming Frog 24.3 on a dedicated server, list mode with JavaScript rendering enabled, 898 URLs: all 49 core pages, all 349 blog URLs, and a random sample of 500 from the 19,345 /project/* pages — 2.6%, drawn with random.seed(20260817) for reproducibility. Proportions from that sample are representative; absolute figures for the UGC section are extrapolations and are labelled as such. We sampled randomly rather than taking the first or most prominent URLs, because prominent pages are systematically biased — they are where fixes get applied first, so testing a rule there reports "works" when the site-wide answer is different.

Site size is measured from the sitemap, not the crawl. The site is 19,743 URLs (19,345 UGC + 349 blog + 49 core), not 898. Page counts derived from a crawl or from pagination are false derivatives: crawl scope is a decision we made, not a property of the site.

Search volume and bids come from Google Ads Keyword Planner — the primary source, exact rather than modelled, and free. Third-party volume estimates were used for nothing in the final figures. Keyword difficulty is on the DataForSEO scale and is not comparable with Ahrefs figures in earlier reports.

Page speed verdicts come from CrUX field data (real Chrome users, 28 days to 15 Aug 2026). PageSpeed Insights lab data is used only to identify causes. Where the two diverge by a factor of several, that is normal and the field data governs. Per-URL field data exists for only 5 URLs because CrUX withholds low-traffic records — a property of the source, not a measurement failure.

The HTTP Last-Modified header was deliberately excluded. It is populated on 397 pages but holds just 8 distinct values, all inside a 31-second window matching our crawl time — it reports CDN cache age. Using it would have marked 397 pages as fresh and erased the freshness findings entirely.

Embeddings are TF-IDF with LSA over 200 components (explained variance 0.634), not neural embeddings — no external API was involved. Cosine similarity was computed over 79,003 page pairs across the 398 core and blog pages; only 8 exceeded 0.85. The UGC class was handled as a class rather than as 124,750 pairs. Boilerplate above 30% document frequency was stripped, and 75 release-note articles plus 23 pagination pages were excluded as template repetition — without that filter, one keyword alone would have produced a false 22-page cannibalisation cluster.

Ahrefs was not used in this run by agreement, so there is no Brand Radar share of voice, no Ahrefs difficulty or content gap, and no cross-check of index completeness. Backlink data comes from DataForSEO, whose index trails Ahrefs by 5–29% on live domains — so every absolute link figure in section 7 is a lower bound. Total third-party data spend for the audit was $2.58.

Corrections we made to our own work, disclosed because they changed conclusions. Our opening hypothesis that robots.txt was blocking the blog was disproved at the primary source and never reached this report. An automated fan-out pass scoring 66% was discarded for 53% after we read all 120 evidence strings. A 90,500-a-month "gap" was withdrawn once we checked that /blog/cat-tools already ranks eighth for it. A 33,100-a-month "opportunity" was withdrawn on bid evidence. Two extractability scores that disagreed were reconciled by rereading the pages rather than averaged. And a Wikipedia "gap" was rejected after reading the article in full. In each case the derived signal was wrong and the primary source corrected it.

B. Blocked checks and what closes them

Eight checks were not performed. None of them is reported as a zero result.

What we needFrom whomWhat it unblocks
GA4 Viewer access for [email protected] plus the numeric Property IDclientSix metrics at once: LLM traffic share, share of organic, engagement rate, bounce rate, ARPU, and top LLM landing pages. ARPU additionally needs e-commerce or conversion events configured
Search Console access (Full or Restricted) on the same emailclientCTR gap and pages-decay — two central checks of section 4 — plus the long tail of queries earning impressions without a ranking position
Ahrefs Brand Radaragency toolingQuantitative share of voice. We will not substitute another tool and publish the result under that name
Reddit API accessagency toolingCommunity mentions across 15 subreddits. Currently 403 from our environment — not "no mentions"
Contractor link inventory (one file)agencyConverts 645 link-gap domains into true Pitch Targets. Pure subtraction — no new data spend
Server or CDN logs, 30–90 daysclientExact LLM bot crawl activity. Optional — proxy signals already cover this section
Current price list, or access to the pricing endpointclientConfirms whether the marked-up 0/50/150/450 tiers are correct before they are published as text
ICP description and the author listclientSharpens case-study gap analysis and the author-page build
ISO 27001 certificate number and issuing bodyclientLets the held certification be published as verifiable markup
Confirmation of the founding yearclientMarkup says 2008; /about, Wikipedia, Wikidata and Crunchbase all say 2009. Four sources to one, including your own site — we recommend 2009 but will not change it without confirmation

C. Data artifacts

This report is a thin layer over the data. Every finding is traceable to a row in the client spreadsheet, and the tabs below are linked individually from the sections above.

TabRowsContents
fixes77Prioritised fix list across all sections
insights614Every finding with severity and a ready action
pages898Full crawl with per-page scores
schema-errors1,599Structured-data defects by page and field
keyword-gaps411Keyword gap pool with bids and competitor positions
do-not-replicate890Excluded donor domains, with a reason each
link-gap-outreach645Outreach targets linking to competitors
fanout120Every sub-question check with its evidence string
existing-pages-gap89Keywords mapped to pages that already exist
inventory-buylist44Donors with a paid placement route
authors42All authors with article counts
cannibalization40Keyword and semantic overlap clusters
case-studies39Case studies with ICP segment and quality flags
clusters17Keyword clusters with volume, difficulty and bid
profiles17Review and directory profiles
authority-placements15Platforms assessed for placement
page-plan15Proposed pages with targets and priority
noindex-in-sitemap13URLs submitted for indexing while forbidding it
top1010Money pages with all three audit levels scored
interlinking7Orphan money pages and missed link opportunities

▧ Open the full audit spreadsheet

Prepared by ESA Digital · Aug 17, 2026 · Data sources: Screaming Frog crawl (898 URLs) · Google Keyword Planner · DataForSEO Labs & LLM Mentions · CrUX & PageSpeed Insights · Common Crawl · live bot checks · 12 ChatGPT spot-checks
This report contains client-confidential data.