YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

gemini-ucd-80k

The model maps Gemini text embeddings to 80,513 labeled concept activations. Its float32 encoder selects the top 128 scores among those concepts, producing at most 128 nonzero activations per text.

Input

Provide a NumPy array of shape (N, 3072) containing embeddings from gemini-embedding-2-preview, using the RETRIEVAL_DOCUMENT task type. Do not apply additional unit normalization. Generate the embeddings separately; this package takes embeddings as input rather than raw text.

Usage

Download the model files and run the following from the model directory:

pip install -r requirements.txt
import numpy as np
from gemini_ucd import GeminiUCD

model = GeminiUCD.from_pretrained(".", device="cpu")
embeddings = np.load("embeddings.npy", mmap_mode="r")
activations = model.encode(embeddings, batch_size=64)

print(activations.shape)  # (N, 80513), a SciPy CSR sparse matrix
print(model.explain(activations[0], top_n=10))

model.explain(row, top_n=10) returns up to 10 active concepts from one sparse output row, ordered by decreasing activation. Each result is a dictionary with concept_id, label, and activation. It looks up the saved labels; it does not run another model or make an API call. top_n controls how many concepts are displayed, not the encoder's top-128 selection.

Use device="cuda" for an NVIDIA GPU. Keep the model in float32. For large datasets, model.encode_iter(embeddings, batch_size=64) yields sparse batches without collecting the full output in memory.

Example

With model loaded above, embed a sentence and display its 50 highest-activating concepts. The embedding step uses the Google GenAI SDK and requires your Gemini API key in the GEMINI_API_KEY environment variable:

pip install google-genai
import os
from google import genai
from google.genai import types

text = "Djokovic Overwhelms Nadal for the Wimbledon Title"
client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
response = client.models.embed_content(
    model="gemini-embedding-2-preview",
    contents=[text],
    config=types.EmbedContentConfig(
        task_type="RETRIEVAL_DOCUMENT",
        output_dimensionality=3072,
    ),
)
example_embeddings = np.array(
    [embedding.values for embedding in response.embeddings], dtype=np.float32
)
example_activations = model.encode(example_embeddings)

for concept in model.explain(example_activations[0], top_n=50):
    print(
        f"{concept['concept_id']:5d}  "
        f"{concept['activation']:.6f}  {concept['label']}"
    )

The output lists the concept ID, activation score, and saved label for each of the top 50 active concepts, ordered from highest to lowest activation.

Concepts and prevalences

  • labels.csv has one row per output coordinate: concept_id is the column index (0–80512), label is the concept name, and discussion is a description of activating examples used to produce the label.
  • prevalences.csv and prevalences.npz contain the same heuristic corpus prevalences at activation thresholds >0.05 and >0.1, with counts and one-sided 99% upper confidence bounds. These can be used to identify concepts that are enriched in your dataset.
Downloads last month
68
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support