knot2vec

Knot2Vec logo: the torus knot T(3,5) in rainbow colours on a dark tile

Embeddings for knot diagrams. Two diagrams of the same knot, however tangled, should land close together; diagrams of different knots should land apart. Part of Weird2Vec, embedding models for data nobody embeds.

What is a knot diagram? A knot is a closed loop in space. Drawn on paper it becomes a diagram: a curve with crossings, each marked over or under. The same knot has infinitely many diagrams, related by Reidemeister moves. Telling whether two diagrams show the same knot is hard; this model learns a fast, approximate answer.

Two tangled diagrams with 54 and 26 crossings, each named correctly by Knot2Vec as the torus knot T(3,5) and the Conway knot

Two tangles the model names correctly. Results below say how often that happens. Try your own in the demo.

🚀 Usage

The repo holds the script that trained the model, so one command embeds a PD code and names the nearest catalogued knots:

uv run https://huggingface.co/jgalego/knot2vec/resolve/main/knot2vec.py embed --pd "[[1,5,2,4],[3,1,4,6],[5,3,6,2]]"

It prints the five nearest of the 12,965 prime knots up to 13 crossings, guesses for signature, determinant and hyperbolic volume, and the 256-dimensional embedding.

🧶 Data

jgalego/knot2vec-diagrams: every prime knot up to 13 crossings in KnotInfo, with invariants, and diagrams of each knot scrambled by random Reidemeister moves in SnapPy. 10% of the knots are held out of training entirely.

🏋️ Training

A diagram becomes its Gauss sequence: walk the knot and, at each crossing, record whether the strand passes over or under and the crossing's sign. Crossings are named by first visit and the walk starts at a random edge. A transformer encodes the sequence, and a contrastive loss pulls two scrambles of the same knot together against the rest of the batch. Small heads predict signature, determinant and volume.

Parameters 4,974,095 (6 layers, width 256)
Steps 10000, 512 knots (two diagrams each) per step
Learning rate 0.0001, cosine; temperature 0.05
Final loss 1.1689 (contrastive 0.5045)
Hardware NVIDIA A10G, 92 min

📊 Results

Each test diagram is a fresh scramble. The model names it by the nearest canonical KnotInfo diagram among all 12965 knots. Seen knots were in training, as other diagrams; unseen knots never were.

Knots n Top-1 Top-5 Signature Determinant ±10% Volume MAE
seen 46676 0.874 0.978 0.926 0.264 1.21
unseen 5184 0.867 0.976 0.896 0.246 1.373

Chance top-1 is 7.7e-05. For an exact answer, use SnapPy's identify().

Every catalogued knot's embedding flattened with UMAP, coloured by signature

Every catalogued knot's embedding, flattened with UMAP and coloured by signature. The map is ordered by signature, from positive to strongly negative.

⚠️ Limitations

  • Prime knots up to 13 crossings only, and only the chirality KnotInfo lists; a mirror image is a different input.
  • Use it to shortlist candidates, then confirm with an exact tool such as SnapPy.
  • Diagrams over 64 crossings are out of range.

📚 Related work

The closest work is Halverson & Ruehle (2025), who train contrastive and generative models to embed braid words of the same knot at the same point. Knot2Vec works on PD codes scrambled by Reidemeister moves instead, covers every prime knot up to 13 crossings, and holds out 10% of the knots.

Articles

Data and software

  • KnotInfo: table of knot invariants by C. Livingston and A. H. Moore, and its Python package database_knotinfo.
  • SnapPy: M. Culler, N. M. Dunfield, M. Goerner and J. R. Weeks' program for the topology and geometry of 3-manifolds, with the spherogram link library.

Blogs and press

Downloads last month
114
Safetensors
Model size
4.97M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train jgalego/knot2vec

Space using jgalego/knot2vec 1

Papers for jgalego/knot2vec