YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
LBind: Multi-Modal Embeddings
Model Details
Model Description
LBind is a multi-modal embedding model that supports image, video, audio, text, and 3D point cloud inputs. All modalities are projected into a shared embedding space, enabling cross-modal similarity computations. The model builds on top of three other models; Perception Encoder, ImageBind, and Uni3D. As indicated by the figure in the top, data is first embedded individually by the three said models. Audio and 3D point cloud embeddings are successively projected with an MLP into the embedding space of the Perception Encoder. The model produces unit-norm embeddings directly usable for similarity comparisons via dot-products ([cosine similarity]).
This version loads all encoders. If you do not need all modalities, please refer to the audio-vision and 3D-points-vision only models.
Uses
Direct Use
The model is intended to be used with direct file-inputs of the said modalities; image, video, audio, 3D, and text. It will produce a 1024 dimension embedding per input, suited for similarity computations.
Downstream Use
The model could be used to build multimodal LLMs, generative models, and systems that perceive their surroundings via both visual, audio, and point cloud embeddings.
Bias, Risks, and Limitations
The model was built on data specified in the paper. As such, it will be biased towards data that "lives on the internet." For specific use-cases, a subsequent fine-tuning stage may be necessary.
How to Get Started with the Model
Option 1
If you want to work within the repository, use uv to install the necessary dependencies.
git clone https://github.com/strictnullchecks/lbind
cd lbind
uv sync
Option 2
You can also install it as an external dependency for another project:
# Option 2.a
python -m pip install git@https://github.com/strictnullchecks/lbind
# Option 2.b; or install a local, editable version
git clone https://github.com/strictnullchecks/lbind
cd /path/to/your/project
python -m pip install -e /path/to/lbind
If you are running a project with pytorch=2.8.0, you should install torchcodec=0.7.0 (as opposed to the =0.8.0) which is automatically installed with uv. torchcodec=0.8.* matches pytorch=2.9.0.
The 3D point cloud backbone has a few custom CUDA kernels that you might want to compile. To do that, you will have to do use Option 1 or Option 2.b above to get a local copy of the repository and compile the kernels.
Loading the Model
import torch
from lbind import LBindModel, LBindProcessor
model = LBindModel.from_pretrained("strictnullchecks/lbind-full")
processor = LBindProcessor.from_pretrained("strictnullchecks/lbind-full")
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = model.to(device).eval()
processor = processor.to(device)
Processing Multi-Modal Inputs
inputs = {
"image": ["examples/dog.png", "examples/cat.png"],
"video": ["examples/dog.mp4", "examples/cat.mp4"],
"audio": ["examples/dog.mp4", "examples/cat.mp4"],
"text": ["A dog is howling in the street", "A cat is sleeping on the couch"],
"points": ["examples/dog_point_cloud.npy", "examples/cat_point_cloud.npy"],
}
with torch.inference_mode():
batch = processor(inputs, return_tensors="pt") # set text_file_paths=True if passing text file paths instead of strings
outputs = model.forward(**batch)
Computing Cross-Modal Similarities
keys = list(outputs.keys())
for i, modality in enumerate(keys):
for j, modality2 in enumerate(keys[i + 1:]):
result = outputs[modality] @ outputs[modality2].T
print(f"{modality} x {modality2}:")
print(result.cpu().detach().numpy())
print('='*26)
Expected Output:
image x video similarity:
[[0.48 0.42]
[0.41 0.6 ]]
==========================
image x audio similarity:
[[0.07 0.05]
[0.02 0.12]]
==========================
image x text similarity:
[[0.16 0.07]
[0.08 0.14]]
==========================
image x points similarity:
[[0.2 0.19]
[0.18 0.19]]
==========================
video x audio similarity:
[[0.19 0.08]
[0.03 0.16]]
==========================
video x text similarity:
[[0.26 0.05]
[0.11 0.14]]
==========================
video x points similarity:
[[0.24 0.15]
[0.17 0.26]]
==========================
audio x text similarity:
[[ 0.12 -0. ]
[ 0.07 0.09]]
==========================
audio x points similarity:
[[0.13 0.06]
[0.1 0.12]]
==========================
text x points similarity:
[[0.19 0.14]
[0.05 0.18]]
==========================
Note: The image/video similarity is significantly higher because they share the same vision encoder.
Compile PointNet2 CUDA ops (optional)
If you have CUDA available, consider building the PointNet2 custom ops used for embedding point clouds to get faster inference:
cd src/lbind/models/uni3d/pointnet2_ops && \
uv run python -c "import torch,sys; sys.exit(0 if torch.cuda.is_available() else 1)" && \
MAX_JOBS=$(nproc) uv run python setup.py build_ext --inplace
We have modified the code slightly in
src/lbind/models/uni3d/pointnet2_ops/pointnet2_utils.pyto have a fallback torch implementation in order for the model to be executable on no-GPU hardware.
Evaluation
We have evaluated the model on multiple benchmarks. We highlight that LBind is performing close to as well as models 4 and 17 times larger.
Figure 1: An average of the 13 benchmarks presented in the two tables below, plotted against model size.
- Downloads last month
- 5

