Food Label Scanner

The AI That Reads and Understands Every Product

An AI engine that reads product fronts, barcodes, and ingredient lists — then indexes them, organizes them, assigns a unique ID to each product, and recognizes what they actually are. One messy shelf of packaging becomes clean, structured, machine-readable product data.

OVERVIEW

Read the front, the barcode,
and the ingredient list


Point a camera at any product and the scanner reads all three layers of information at once: the marketing name on the front, the barcode, and the fine print of the ingredient list. It then turns that raw text into a structured, indexed record — parsing the exact structure of the ingredient statement, resolving each token to a shared catalogue of canonical ingredients, and assigning every distinct product a stable, unique ID. No manual data entry, no template setup. One messy shelf of packaging becomes clean, machine-readable product data.

The scanner's operator console reading a product label and resolving ingredients to canonical entries and their aliases
A real product processed end-to-end: the scanner reads the front and the ingredient label (OCR, right), parses the ingredient statement into a structured list (bottom-left), and resolves each token to a canonical entry — here さば (mackerel), matched against its aliases サバ・鯖・さば節・さば削り節… so every spelling and form points to one product identity.
82,816
products analyzed
across 3 domains
~98%
of ingredient
occurrences resolved
~6,000
canonical ingredients
in the shared catalogue
28
Japanese statutory
allergen groups

CHALLENGE

The same product,
written a hundred different ways


One Product, Many Names

The same item appears as kanji, katakana, hiragana, or an abbreviation depending on who wrote the label. Systems treat them as different products.

Unstructured Ingredient Lists

Ingredient lists are dense, free-form text. Extracting and comparing them by hand is slow and error-prone.

No Shared Product ID

Without a stable identity per product, catalogs, inventories, and analyses can't be reliably linked or deduplicated.

CAPABILITIES

An indexing engine for the physical shelf


Multi-Layer Reading

Reads the product front, barcode, and ingredient list in a single pass, capturing both what a product is called and what it contains.

Indexing & Organization

Parses each ingredient list into an ordered, structured index and organizes products into a clean, queryable catalog.

Unique Product IDs

Assigns every distinct product a stable individual ID, so the same item is recognized consistently across sources and over time.

Semantic Recognition

Understands the meaning behind the text — recognizing that different labels, spellings, and forms can all refer to the same underlying product.

PRODUCT ALIASING

How the scanner knows two labels mean one product


1

Aliasing by Writing

Japanese products can be written in kanji, katakana, or hiragana — often for the exact same word. The scanner normalizes across writing systems and spellings so those variants resolve to a single product.

生姜 · しょうが · ショウガ → ginger
2

Aliasing by Semantic Understanding

Beyond spelling, the engine understands form and preparation. Powdered, sliced, grated, or whole — it recognizes they are all the same core ingredient and links them to one identity.

ginger · powdered ginger · sliced ginger → ginger
3

Aliasing by Location

Place names on labels — regions, origins, and production areas — are recognized and understood, so origin-qualified products are grouped correctly instead of fragmenting the catalog.

Kochi ginger · domestic ginger → recognized origin of ginger

RECIPE UNDERSTANDING

Decoding composite ingredients — and translating them


Recipe Decomposition

Manufacturers often list a prepared component by name — "egg salad" — without spelling out what's in it. The engine decodes the recipe into its constituents, so hidden ingredients and allergens surface.

egg salad → egg + mayonnaise (+ vinegar, sugar…)

The Label Always Wins

When the manufacturer does spell out the breakdown in parentheses, that explicit list overrides the inferred recipe. The printed label is always the source of truth.

egg salad (egg, vegetable oil, brewing vinegar…) → uses the listed ingredients

Cross-Language Resolution

Some Japanese ingredients have no direct foreign name. The engine resolves them to the closest equivalent — making it a powerful translation aid for anyone reading a label in another language.

かまぼこ (kamaboko) → "fish cake (surimi)"

USE CASE

In the field: the Labemiru app


Labemiru logo

Labemiru — understand what's inside a product, at a glance.

Labemiru is a consumer app built on the Food Label Scanner. A shopper points their phone at a product — say, a bag of ginger candy — and the scanner reads the front, barcode, and ingredient list at once, then explains the contents in plain language.

  • Reads "生姜あめ" on the front and "しょうが" in the ingredient list — and recognizes both as the same ginger.
  • Maps allergens to Japan's 28 statutory groups, including those hidden inside compound ingredients, and grades additive concern.
  • Re-reads the same label through lifestyle lenses — vegan / vegetarian, pork & halal awareness, pregnancy caution, pet safety, and an MSG marker.

Behind the app, the same engine spans three domains — food, personal care, and pet food — and is available through a lightweight barcode → analysis API. It is an informational tool, not a medical device.

Visit labemiru.com

The Labemiru app comparing two food products side by side, each scored with its ingredient breakdown

Turn packaging into structured product data.

Interested in the Food Label Scanner or the Labemiru app built on it?
Tell us about your products and use case — we're happy to help.

Request Materials / Contact