Public methodology

How the Hindi font conversion pipeline works

SahayakTools combines exact legacy-mapping boundaries, positional Devanagari rules, rich-document handling and fail-closed review checks. This page explains what the system does—and what it deliberately refuses to guess.

Pipeline

Six stages from mapping confirmation to safe export

1. Confirm the exact mapping

A route selects one evidence-backed legacy mapping, not every font with a similar family name. Kruti Dev 010, DevLys 010, standard Chanakya, Shree714 and the supported AMS reference encoding have separate boundaries. Source conversion stays guarded when the exact variant cannot be inferred safely.

2. Sanitize rich clipboard content

When content is pasted from Word or a browser, unsafe executable elements and event attributes are removed. Supported paragraph, emphasis, list and table elements are retained, while safe Word typography such as font size, line spacing and paragraph geometry is inlined before style blocks are discarded.

3. Convert complete font runs

Legacy conversion is not treated as independent character replacement. Compatible text split across inline spans is regrouped within block boundaries so matras, conjuncts and reph rules can be applied to the semantic sequence. Non-source font runs are left unchanged.

4. Protect only tokens that are actually safe

URLs, email addresses, filenames, paths, dates and similar high-confidence document tokens can be preserved in unscoped plain-text workflows. Inside a confirmed legacy-font run, ASCII symbols may themselves be glyph codes—so characters such as @, # or / are not blindly protected from the selected mapping.

5. Check semantics, not just round trips

Golden reference vectors, exact font-layout evidence, generated cases and mixed-document invariants are used together. A round trip is only one signal: two wrong inverse tables can round-trip perfectly, so round-trip consistency alone is never treated as proof that a legacy mapping is correct.

6. Preserve font-run meaning on export

Unicode output can be copied as normal text. Reverse legacy output may need separate font runs for encoded Hindi and preserved English, numbers or symbols. When those runs are part of the meaning, Raw text, TXT and plain Share fail closed and Copy for Word is required instead of silently dropping formatting.

Why sequence matters

Legacy Hindi is an encoding problem, not a font-style problem

In Unicode, a Devanagari character has a standard code point. In older font encodings, Latin-code positions and custom glyph sequences were used to display Hindi. Merely applying Mangal to a Kruti Dev string therefore exposes the underlying Latin-looking codes instead of converting the text.

Some visible forms also depend on order. The short-i matra, reph and conjunct sequences may be stored differently from their visual position. That is why conversion needs ordered replacements and positional handling rather than a one-character lookup alone.

Known limits

What the converter will not pretend to know

!Plain ASCII-looking legacy text usually cannot prove its exact font/version by itself; confirm the source mapping or preserve Word font metadata.
!Kruti Dev 020/030, DevLys 020/030, Walkman-Chanakya-905/901, ShreeLipi715, Dev Ratna Universal, Shree Dev 0702 and unrelated AMS variants are not assumed interchangeable with the mappings currently supported.
!Decorative glyphs, private-use characters, equations and unusual symbols can be unsupported. The converter should preserve or flag uncertainty rather than silently invent a mapping.
!Changing from a legacy font to a Unicode font can change glyph widths even at the same point size. SahayakTools preserves source typography and table structure where available, but pixel-identical wrapping across different fonts is not guaranteed.
!Native DOCX conversion remains hidden until its Office regression and formatting-preservation suite is executed in a real non-Actions environment. The public rich workflow uses Copy for Word rather than a fake HTML file renamed as .doc.
!A high round-trip consistency score shows internal mapping stability only; it is not semantic, legal or filing correctness.

What the release gate is designed to test

The converter suite combines primary/reference glyph vectors, reverse semantic checks, generated consonant/matra/conjunct cases, protected-token stress tests, Word span fragmentation, exact font-family boundaries, table/typography preservation and mixed Hindi/English/number/punctuation cases across the registered converters. A converter is not release-ready merely because its own reverse table round-trips.

View evidence examples

Privacy model

The conversion and quality-analysis functions execute in the web page. Optional history is off by default; when enabled, recent entries are written to localStorage in the same browser. Users can clear that history from the studio.

Read privacy policy