Formula image to typst model TypLens-V1 released — avaliable on TypstPad

Hi everyone!

I’m happy to share that TypLens-V1 has been released and is now available to try at typstpad.com!

TypLens is a lightweight image-to-Typst model made to run in browsers that turns formulas into Typst code. When you find a formula in a paper or PDF, you can use a screenshot as a starting point instead of retyping the entire expression.

Examples (All results from typstpad.com)

cal( L ) _ ( upright( D P O ) ) ( pi _ ( theta ) ; pi _ ( upright( r e f ) ) ) = - bb( E ) _ ( ( x , y _ ( w ) , y _ ( l ) ) tilde cal( D ) ) lr( [ log sigma lr( ( beta log frac( pi _ ( theta ) ( y _ ( w ) divides x ) , pi _ ( upright( r e f ) ) ( y _ ( w ) divides x ) ) - beta log frac( pi _ ( theta ) ( y _ ( l ) divides x ) , pi _ ( upright( r e f ) ) ( y _ ( l ) divides x ) ) ) ) ] )

bold( mu ) _ ( theta ) ( bold( upright( x ) ) _ ( t ) , t ) = tilde( bold( mu ) ) _ ( t ) lr( ( bold( upright( x ) ) _ ( t ) , frac( 1 , sqrt( macron( alpha ) _ ( t ) ) ) lr( ( bold( upright( x ) ) _ ( t ) - sqrt( 1 - macron( alpha ) _ ( t ) ) thin bold( epsilon.alt ) _ ( theta ) ( bold( upright( x ) ) _ ( t ) ) ) ) ) ) = frac( 1 , sqrt( alpha _ ( t ) ) ) lr( ( bold( upright( x ) ) _ ( t ) - frac( beta _ ( t ) , sqrt( 1 - macron( alpha ) _ ( t ) ) ) thin bold( epsilon.alt ) _ ( theta ) ( bold( upright( x ) ) _ ( t ) , t ) ) )

mat( delim: #none , align: #center , L ( gamma , phi.alt ; alpha , beta ) = zws , log Gamma lr( ( display( sum _ ( j = 1 ) ^ ( k ) alpha _ ( j ) ) ) ) - display( sum _ ( k = 1 ) ^ ( k ) log Gamma ( alpha _ ( i ) ) ) zws ; zws , quad + display( sum _ ( k = 1 ) ^ ( k ) alpha _ ( i ) - 1 ) ) lr( [ Psi ( gamma _ ( i ) ) - Psi lr( ( display( sum _ ( j = 1 ) ^ ( k ) gamma _ ( j ) ) ) ) ] ) zws ; zws , quad + display( sum _ ( k = 1 ) ^ ( N ) sum _ ( k = 1 ) ^ ( k ) delta _ ( m ) lr( [ Psi ( gamma _ ( i ) ) - Psi lr( ( display( sum _ ( j = 1 ) ^ ( k ) gamma _ ( j ) ) ) ) ] ) ) zws ; zws , quad + display( sum _ ( n = 1 ) ^ ( N ) sum _ ( l = 1 ) ^ ( k ) sum _ ( l = 1 ) ^ ( N ) delta _ ( m ) w _ ( n ) ^ ( l ) log beta _ ( l j ) ) zws ; zws , quad - log Gamma lr( ( display( sum _ ( j = 1 ) ^ ( k ) gamma _ ( j ) ) ) ) + display( sum _ ( k = 1 ) ^ ( k ) log _ ( n ) Gamma ( gamma _ ( l ) ) ) zws ; zws , quad - display( sum _ ( k = 1 ) ^ ( k ) ( gamma _ ( l ) - 1 ) ) lr( [ Psi ( gamma _ ( l ) ) - Psi lr( ( display( sum _ ( j = 1 ) ^ ( k ) gamma _ ( j ) ) ) ) ] ) zws ; zws , quad - display( sum _ ( m = 1 ) ^ ( N ) sum _ ( l = 1 ) ^ ( k ) delta _ ( m i ) log phi.alt _ ( m i ) . ) )

Oops! This is a long formula and the model is making some mistakes.

Highlights

  • Browser-compatible inference with a relatively small, 29.4-million-parameter model.
  • MIT open-sourced, downloadable model weights in both full-precision FP32 (~118 MB) and compact INT8 (~34 MB) ONNX variants.

This is still an experimental release. Long expressions, subscripts, and subtle symbol or style differences can still cause mistakes, so please check the generated output, even when it compiles successfully.

Try it: https://typstpad.com
GitHub: GitHub - dbccccccc/TypLens · GitHub
Huggingface: dbcccc/TypLens · Hugging Face

I’d love to hear how it works with your formulas! Examples where recognition fails would be especially helpful—please share the formula image, the generated output, and the expected result here or in a GitHub issue.

Thanks for giving it a try!

8 Likes

Nice! It works remarkably well, even on equations found in in older papers. I did find some bugs, for example this image:


produces a mess:

frak( B ) = sum _ ( script( mat( delim: #none , align: #center , a < A zws ; b < B ) ) ) chevron.l f \, g chevron.r = sum _ ( script( mat( delim: #none , align: #center , a < A ) ) zws ; zws , zws , zws , zws , zws , zws ; zws , zws , zws , zws , zws , zws , zws , zws , zws ; zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , zws , B ) ) ) ) ) ,

image

Less egregiously, I found that it sometimes doesn’t recognise alternative styles or overlines correctly (particularly in older papers), in these images
image


$frak(B)$ is recognised as $bold( upright( B ) )$ and in $overline(a)/b - overline(b)/a$ the overlines are recognised as tildes.

There are also some weird quirks where 2 or more lines under e.g. a sum become $sum_mat(delim: #none, align: #center, line1; line2)$ instead of the simpler $sum_(line1 \ line2)$ and fractions become $frac(A, B)$ instead of $A / B$ but these don’t affect the visual output

Other than that, great job!

1 Like

Hi aarnent, thanks for the detailed examples!

I tried the first formula with different crops and found something interesting: the model can recognize it correctly in one image, but cropping it slightly tighter changes the result.

With this particular crop, I got the correct result:

frak( B ) = sum _ ( script( mat( delim: #none , align: #center , a < A zws ; b < B ) ) ) chevron.l f , g chevron.r = sum _ ( script( mat( delim: #none , align: #center , a < A zws ; b < B ) ) ) integral _ ( Gamma backslash cal( H ) ) f ( z ) overline( g ( z ) ) y ^ ( - 2 ) d x d y

See output in TypstPad

So it can handle this expression, but its recognition isn’t stable enough across small changes to the input. I haven’t isolated the exact cause yet, but based on my testing so far, both the cropping and clarity will affect the output.

TypLens currently resizes the entire input image to 384 × 384 using bicubic interpolation, without preserving the original aspect ratio. This means that even cropping away a little whitespace can change the formula’s scale, proportions, and position in the processed image, while resizing can affect fine details such as overlines. This may explain why one crop is recognized correctly while a slightly tighter crop produces errors, though I’m still investigating the exact cause.

For the mat(...) and frac(...) cases, I agree that simpler output would be easier to read and edit. I haven’t found a suitable ready-made image-to-Typst dataset, so I use image-to-LaTeX datasets and convert their annotations to Typst for training. That conversion happens during data preparation; the model itself generates Typst directly, without an intermediate LaTeX step. The converter I use produces forms such as frac(...) , which then appear in the training targets. I’d like to make those targets more consistent and favor simpler Typst syntax in future versions. The repository has more details about the training-data sources.

Thanks again for testing it on older papers and sharing where it falls short! This is exactly the kind of feedback that helps me understand what needs improving beyond my own test examples.

Happy to help! I figured some kind of latex to typst conversion was going on at some point, these “non-canonical” representations are something i’ve experienced when using pandoc. I suppose this conversion step is somewhat unavoidable since as far as I could tell most of these LaTeX datasets use “real” formulas from arXiv papers (which are sadly not in typst yet) etc, rather than building some sort of formula generator.

Hi! Since the last release, I’ve done more testing and training, and TypLens V1.1 is now available on GitHub and Hugging Face. You can try it at typstpad.com.

I’ve revised the image preprocessing and retrained the model to improve recognition stability across different input images. More details are available in the code and documentation on GitHub. In my tests, the new version correctly handles all three examples you reported and also performs better on my local test suite.

For future versions, I’m planning to improve support for longer formulas and make the output more typstic. Thanks again for the helpful feedback!

2 Likes