Skip to content
Dataset By Humans
On requestPhotos· Southeast Asia

Vietnamese Scene Text Image Dataset for OCR

Real shop signs, menus, banners and street notices in Vietnamese, full of diacritics, mixed fonts and imperfect lighting.

On request: this dataset is produced to your specification after you order. Typically 1–2 weeks for 500 images and 3–5 weeks for 2,000; transcriptions add time.

  • Bilingual parking sign reading 'Toà/Building BS11 BS12' with an EV charging station sign, in front of high-rise apartments
  • Night street view of a print and photocopy shop with large Vietnamese signs reading 'Nhà in - Photocopy Nam Dũng' and parked motorbikes
  • Shopfront of a home-appliance company with dense Vietnamese signage, neighbouring food shop signs and motorbikes on the pavement
  • Corner snack store with large red and white Vietnamese signs, slogans and a hotline number, with street traffic below

Pricing

Fixed prices for standard packages, produced to your spec. Pick a package and a license to see the exact total, or ask for a custom quote.

Packages compared
FeatureStarterProCustom
images5002,000Any quantity
What's included
  • 500 full-resolution images
  • metadata.jsonl with device, exposure and topic labels
  • Datasheet, license and checksums
  • 2,000 full-resolution images
  • Everything in Starter
  • Human-written English caption per image
Your quantity, annotations and deadline
Datasheet, license, checksumsIncludedIncludedIncluded
Research license$490$0.98 / image$1,490$0.75 / imageQuote
Commercial license$1,250$2.50 / image$3,900$1.95 / imageQuote
Exclusive license$4,900$9.80 / image$14,900$7.45 / imageQuote

How ordering works

  1. 1Choose a package and license and place your order. No payment is taken online.
  2. 2Within 48 hours our team contacts you to confirm the spec and invoices a 50% deposit.
  3. 3We produce and deliver in parts so you can review early. Typically 1–2 weeks for 500 images and 3–5 weeks for 2,000; transcriptions add time.
  4. 4Pay the balance and download the final dataset.

Custom quote

Need another quantity, an Enterprise license or annotations for this dataset? Available here:

  • Topic labels
  • Full-image transcriptions
  • Word-level boxes with transcription
Request a custom quote

Overview

Vietnamese uses the Latin alphabet with stacked diacritics, so a single character can carry two marks. OCR models trained mostly on English text routinely drop or confuse these marks, especially on hand-painted signs, curved banners and low-light menus.

This dataset contains photos of Vietnamese text as it appears in daily life: storefront signs, restaurant menus, price boards, banners, notices and packaging, photographed at natural angles with real-world blur, glare and occlusion.

Sample photos

Real photos from this topic, shown at reduced resolution. Full-resolution samples are emailed on request.

  • Bilingual parking sign reading 'Toà/Building BS11 BS12' with an EV charging station sign, in front of high-rise apartments
  • Night street view of a print and photocopy shop with large Vietnamese signs reading 'Nhà in - Photocopy Nam Dũng' and parked motorbikes
  • Shopfront of a home-appliance company with dense Vietnamese signage, neighbouring food shop signs and motorbikes on the pavement
  • Corner snack store with large red and white Vietnamese signs, slogans and a hotline number, with street traffic below

Get a free sample pack

We email you full-resolution samples for this topic.

Used only to send your samples. Privacy policy.

Technical specifications

Capture device
Modern smartphones (device model recorded per image in EXIF)
File format
JPEG (original HEIC/RAW on request)
Resolution
Native sensor resolution, typically 12 MP or higher
Text types
Shop signs, menus, price boards, banners, notices, packaging
Language
Vietnamese, with some mixed Vietnamese–English text
Metadata
metadata.jsonl with device, exposure, timestamp, dimensions, topic labels
Privacy processing
Faces and license plates blurred, GPS removed
License
Research, Commercial, Enterprise or Exclusive

Use cases

  • Scene text recognition

    Fine-tune OCR and text-spotting models (PaddleOCR, TrOCR, docTR) to read Vietnamese diacritics accurately.

  • Multilingual VLM evaluation

    Benchmark how well vision-language models read and translate Vietnamese text in natural images.

  • Translation & accessibility apps

    Train camera-translation and screen-reader features for travellers and visually impaired users.

  • Retail & mapping

    Extract business names and categories from storefront imagery for maps and local search.

How this data is made

  1. 1Captured on a smartphone by a person
  2. 2Duplicates and near-duplicates removed
  3. 3Blurry and broken images filtered out
  4. 4Faces and license plates detected and blurred
  5. 5Every image reviewed by a person
  6. 6GPS stripped; device EXIF kept
  7. 7Labels, captions and metadata.jsonl generated
  8. 8Datasheet and license packaged with the delivery
Read our full process

Provenance & legal

  • 100% made by people: no scraping, no generative AI
  • Datasheet documenting how the data was made, checked and its limitations
  • Commercial license that lets you keep models trained on the data
  • Supports training-data documentation under the EU AI Act
Licensing options

Dataset-specific notes

  • Phone numbers and personal names on signs can be blurred on request; business names are kept.
  • Bystanders' faces are blurred and children are excluded.
  • Text transcriptions are written by native speakers of the language.

Example metadata record

One line of metadata.jsonl per image. Fields vary with the annotation options you choose; values shown are illustrative.

{
  "id": "vst-000318",
  "file": "images/vst-000318.jpg",
  "topic": "scene-text",
  "width": 4000,
  "device": {
    "make": "Samsung",
    "model": "Galaxy S24"
  },
  "height": 3000,
  "region": "Southeast Asia",
  "privacy": {
    "gps_removed": true,
    "faces_blurred": 0,
    "phone_numbers_blurred": 1
  },
  "text_type": "menu",
  "captured_at": "2026-09-20T19:44:02+07:00",
  "transcription": "PHỞ BÒ TÁI 45.000đ | PHỞ GÀ 40.000đ | BÚN CHẢ 50.000đ"
}

Each delivery contains:

vietnamese-scene-text-ocr-v1/
├── images/
├── metadata.jsonl
├── DATASHEET.md
├── LICENSE.pdf
└── checksums.sha256

Frequently asked questions

Are transcriptions included?

Transcriptions are available on request, written by native speakers. Word-level bounding boxes with text can also be quoted.

Can you include handwritten text?

Yes, hand-painted signs and handwritten menus can be collected. Mention this in your request.

Is the text in the images real?

Yes. Every image is a photo of real signage or print. Nothing is rendered or generated.

How does ordering and payment work?

Create a free account and place your order on this page; nothing is charged online. Our team contacts you to confirm the spec and invoices a 50% deposit, with the balance due on final delivery. The deposit is refunded in full if we cannot deliver the agreed spec.

What is the difference between non-exclusive and Exclusive?

With Research or Commercial we may license the images produced for your order to other buyers later. With Exclusive they are never licensed to anyone else or added to our catalog.

Can I order a different quantity or extra annotations?

Yes. Request a custom quote with the quantity, annotations and deadline you need. Prices per record fall as quantity grows.

Why is a bundle so much cheaper per image than a single image?

A bundle is planned and collected as one job, so capture, review and privacy processing cost far less per photo. Single images are priced for buyers who need just a few specific photos.

Need a variation?

Different topics, languages, difficulty, quantity or labels? Describe it and we reply within 48 hours with matching samples and a quote.

Request a quote
Browse all datasets