Skip to content
unlob

Thai web search

ไทยlang=th

No spaces between words; segmentation is handled at index and query time.

Example

"นโยบายสภาพภูมิอากาศ" — climate policybash
# Thai sources only
curl -H "x-api-key: $UNLOB_API_KEY" \
  "https://api.unlob.com/search?q=%E0%B8%99%E0%B9%82%E0%B8%A2%E0%B8%9A%E0%B8%B2%E0%B8%A2%E0%B8%AA%E0%B8%A0%E0%B8%B2%E0%B8%9E%E0%B8%A0%E0%B8%B9%E0%B8%A1%E0%B8%B4%E0%B8%AD%E0%B8%B2%E0%B8%81%E0%B8%B2%E0%B8%A8&lang=th&limit=10"

# The same meaning, sources in any language
curl -H "x-api-key: $UNLOB_API_KEY" \
  "https://api.unlob.com/search?q=%E0%B8%99%E0%B9%82%E0%B8%A2%E0%B8%9A%E0%B8%B2%E0%B8%A2%E0%B8%AA%E0%B8%A0%E0%B8%B2%E0%B8%9E%E0%B8%A0%E0%B8%B9%E0%B8%A1%E0%B8%B4%E0%B8%AD%E0%B8%B2%E0%B8%81%E0%B8%B2%E0%B8%A8&mode=semantic&limit=10"

How multilingual retrieval works here

The conventional approach runs a separate index per language and translates queries between them, which multiplies infrastructure and loses meaning at every boundary. A single multilingual embedder places all 101 supported languages in one shared space, so texts with the same meaning land near each other regardless of the language they are written in.

The practical consequence is that a query in English can retrieve a Thai passage directly, and vice versa — with no translation step, no per-language index, and no additional resident memory. Only semantic and hybrid modes do this; keyword mode matches tokens, and tokens are monolingual.

Thai does not delimit words with spaces, so a word tokeniser has nothing to split on. Character bigrams are the robust answer — they need no language-specific model, degrade gracefully on mixed-script text, and apply identically at index and query time so the two always agree.

Common questions

How do I search Thai web content?

Set lang=th on any search request. You can also leave it unset — the embedding space holds 101 languages at once, so a query in any language retrieves relevant Thai passages directly.

Do I need to translate my query into Thai?

No. Semantic and hybrid modes cross languages natively. Translating would search for the translation rather than the meaning, which is usually worse. Use lang only when the results must be readable to a Thai-speaking user.

How is Thai tokenised?

As character bigrams, because the script does not delimit words with spaces. The same rule applies at index time and query time, so the two always agree — and it needs no language-specific segmentation model.

Does keyword mode work for Thai?

Within Thai it does. What keyword mode cannot do is cross languages — token matching is monolingual by construction. Use semantic or hybrid mode if you want Thai sources from a query in another language.

Search Thai sources

10,000 free requests a month, no card. All 101 languages are available on every plan.