Configuration Parsing Warning:Invalid JSON for config file config.json

mlx-community/map-anything-fp16

float16 MLX conversion of facebook/map-anything (revision a1d87e90) for the mapanything model in mlx-vlm. MapAnything (Meta and CMU, 3DV 2026) is a feed-forward model for metric 3D reconstruction: from one or more images, plus any optional intrinsics, depth maps and camera poses, one forward pass predicts for every view the metric point map, depth, ray directions, camera pose and intrinsics, a confidence map and a non-ambiguity mask.

This is an unofficial derivative. It holds weights and metadata only; the inference code lives in mlx-vlm. The weights are under CC BY-NC 4.0, like the original: noncommercial use only. For commercial use, see mlx-community/map-anything-apache-fp16, the Apache 2.0 release trained on a different data mix.

What is in model.safetensors (2.65 GB, about half the size of the float32 original):

  • 1.13 B float16 parameters: the linear layers (weights and biases) of the DINOv2 ViT-g/14 encoder (its first 24 blocks) and of the 16-layer multi-view transformer;
  • 95 M float32 parameters: the patch embedding, norms, LayerScales and position embeddings, the geometric input encoders, and the DPT, pose and scale heads;
  • keys renamed to the MLX module tree and conv kernels in channel-last layout. config.json is the original configuration plus model_type.

The encoder and transformer run their matmuls and attention in float16 with a float32 residual stream, norms and activations: the reference's mixed precision picks float16 on Apple GPUs. The heads run in float32, as in the reference.

Use with mlx-vlm

Requires an mlx-vlm version that includes the mapanything model (added after 0.7.4).

import mlx.core as mx
from mlx_vlm import load

model, processor = load("mlx-community/map-anything-fp16")

views = processor.load_images("path/to/images")  # a folder or a list of images
predictions = model.infer(views)                  # one dict per view, lazy
mx.eval(predictions)

pred = predictions[0]
pred["pts3d"]         # (1, H, W, 3) metric world points (first view's frame)
pred["depth_z"]       # (1, H, W, 1) metric depth
pred["camera_poses"]  # (1, 4, 4) OpenCV cam2world
pred["intrinsics"]    # (1, 3, 3)
pred["conf"]          # (1, H, W) confidence
pred["mask"]          # (1, H, W, 1) valid pixels

Any view can add calibration, depth and pose, in any combination:

views = processor.preprocess([
    {"img": image0, "intrinsics": K0, "depth_z": depth0, "camera_poses": pose0},
    {"img": image1, "intrinsics": K1},
    {"img": image2, "camera_poses": (quats_xyzw, trans), "is_metric_scale": False},
])
predictions = model.infer(views)

Command line, writing every output to predictions.npz and a colored point cloud to points.ply:

MLX_ENABLE_TF32=0 python -m mlx_vlm.models.mapanything.generate \
    --model mlx-community/map-anything-fp16 --images path/to/images --output out

On Apple M5 and later GPUs, MLX runs float32 matmuls at TF32 precision by default, which makes the float32 heads the largest source of error. Set MLX_ENABLE_TF32=0 before the first MLX matmul to avoid it (the command line does this); it costs a few milliseconds.

Accuracy and speed

Measured against the float32 PyTorch reference on frames of three outdoor videos (4 views at 518x294, MLX_ENABLE_TF32=0). With this float16 model, depth after scale normalization has a median relative error of 5e-5 to 2e-4 and the metric scale is within 1e-4 to 1.2e-3, on par with the reference's own mixed precision on MPS (6e-5 to 3e-4 and 2e-5 to 1.7e-3). In float32, the MLX port matches the reference to ~1e-6 on every output.

Full infer (network and post-processing) on an Apple M5 Max, 518x294 views. Another GPU application was running during these runs, so treat the numbers as relative:

Views MLX, this model PyTorch MPS, mixed precision
1 0.06 s 0.13 s
4 0.23 s 0.47 s
8 0.48 s 1.73 s
16 1.10 s 5.27 s
32 2.89 s 11.3 s

Citation

@inproceedings{keetha2026mapanything,
  title={{MapAnything}: Universal Feed-Forward Metric 3D Reconstruction},
  author={Keetha, Nikhil and M{\"u}ller, Norman and Sch{\"o}nberger, Johannes and Porzi, Lorenzo and Zhang, Yuchen and Fischer, Tobias and Knapitsch, Arno and Zauss, Duncan and Weber, Ethan and Antunes, Nelson and others},
  booktitle={International Conference on 3D Vision (3DV)},
  year={2026},
  organization={IEEE}
}
Downloads last month
-
Safetensors
Model size
1B params
Tensor type
F32
·
F16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/map-anything-fp16

Finetuned
(1)
this model