Instructions to use mlx-community/map-anything-fp16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/map-anything-fp16 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir map-anything-fp16 mlx-community/map-anything-fp16
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Configuration Parsing Warning:Invalid JSON for config file config.json
mlx-community/map-anything-fp16
float16 MLX conversion of facebook/map-anything
(revision a1d87e90) for the mapanything model in
mlx-vlm.
MapAnything (Meta and
CMU, 3DV 2026) is a feed-forward model for metric 3D reconstruction: from one
or more images, plus any optional intrinsics, depth maps and camera poses, one
forward pass predicts for every view the metric point map, depth, ray
directions, camera pose and intrinsics, a confidence map and a
non-ambiguity mask.
This is an unofficial derivative. It holds weights and metadata only; the
inference code lives in mlx-vlm. The weights are under CC BY-NC 4.0, like the original: noncommercial use only. For commercial use, see mlx-community/map-anything-apache-fp16, the Apache 2.0 release trained on a different data mix.
What is in model.safetensors (2.65 GB, about half the size of the float32
original):
- 1.13 B float16 parameters: the linear layers (weights and biases) of the DINOv2 ViT-g/14 encoder (its first 24 blocks) and of the 16-layer multi-view transformer;
- 95 M float32 parameters: the patch embedding, norms, LayerScales and position embeddings, the geometric input encoders, and the DPT, pose and scale heads;
- keys renamed to the MLX module tree and conv kernels in channel-last
layout.
config.jsonis the original configuration plusmodel_type.
The encoder and transformer run their matmuls and attention in float16 with a float32 residual stream, norms and activations: the reference's mixed precision picks float16 on Apple GPUs. The heads run in float32, as in the reference.
Use with mlx-vlm
Requires an mlx-vlm version that includes the mapanything model (added
after 0.7.4).
import mlx.core as mx
from mlx_vlm import load
model, processor = load("mlx-community/map-anything-fp16")
views = processor.load_images("path/to/images") # a folder or a list of images
predictions = model.infer(views) # one dict per view, lazy
mx.eval(predictions)
pred = predictions[0]
pred["pts3d"] # (1, H, W, 3) metric world points (first view's frame)
pred["depth_z"] # (1, H, W, 1) metric depth
pred["camera_poses"] # (1, 4, 4) OpenCV cam2world
pred["intrinsics"] # (1, 3, 3)
pred["conf"] # (1, H, W) confidence
pred["mask"] # (1, H, W, 1) valid pixels
Any view can add calibration, depth and pose, in any combination:
views = processor.preprocess([
{"img": image0, "intrinsics": K0, "depth_z": depth0, "camera_poses": pose0},
{"img": image1, "intrinsics": K1},
{"img": image2, "camera_poses": (quats_xyzw, trans), "is_metric_scale": False},
])
predictions = model.infer(views)
Command line, writing every output to predictions.npz and a colored point
cloud to points.ply:
MLX_ENABLE_TF32=0 python -m mlx_vlm.models.mapanything.generate \
--model mlx-community/map-anything-fp16 --images path/to/images --output out
On Apple M5 and later GPUs, MLX runs float32 matmuls at TF32 precision by
default, which makes the float32 heads the largest source of error. Set
MLX_ENABLE_TF32=0 before the first MLX matmul to avoid it (the command
line does this); it costs a few milliseconds.
Accuracy and speed
Measured against the float32 PyTorch reference on frames of three outdoor videos (4 views at 518x294, MLX_ENABLE_TF32=0). With this float16 model, depth after scale normalization has a median relative error of 5e-5 to 2e-4 and the metric scale is within 1e-4 to 1.2e-3, on par with the reference's own mixed precision on MPS (6e-5 to 3e-4 and 2e-5 to 1.7e-3). In float32,
the MLX port matches the reference to ~1e-6 on every output.
Full infer (network and post-processing) on an Apple M5 Max, 518x294
views. Another GPU application was running during these runs, so treat the
numbers as relative:
| Views | MLX, this model | PyTorch MPS, mixed precision |
|---|---|---|
| 1 | 0.06 s | 0.13 s |
| 4 | 0.23 s | 0.47 s |
| 8 | 0.48 s | 1.73 s |
| 16 | 1.10 s | 5.27 s |
| 32 | 2.89 s | 11.3 s |
Citation
@inproceedings{keetha2026mapanything,
title={{MapAnything}: Universal Feed-Forward Metric 3D Reconstruction},
author={Keetha, Nikhil and M{\"u}ller, Norman and Sch{\"o}nberger, Johannes and Porzi, Lorenzo and Zhang, Yuchen and Fischer, Tobias and Knapitsch, Arno and Zauss, Duncan and Weber, Ethan and Antunes, Nelson and others},
booktitle={International Conference on 3D Vision (3DV)},
year={2026},
organization={IEEE}
}
- Downloads last month
- -
Quantized
Model tree for mlx-community/map-anything-fp16
Base model
facebook/map-anything