YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
gemini-ucd-80k
The model maps Gemini text embeddings to 80,513 labeled concept activations. Its float32 encoder selects the top 128 scores among those concepts, producing at most 128 nonzero activations per text.
Input
Provide a NumPy array of shape (N, 3072) containing embeddings from
gemini-embedding-2-preview, using the RETRIEVAL_DOCUMENT task type.
Do not apply additional unit normalization. Generate the embeddings separately;
this package takes embeddings as input rather than raw text.
Usage
Download the model files and run the following from the model directory:
pip install -r requirements.txt
import numpy as np
from gemini_ucd import GeminiUCD
model = GeminiUCD.from_pretrained(".", device="cpu")
embeddings = np.load("embeddings.npy", mmap_mode="r")
activations = model.encode(embeddings, batch_size=64)
print(activations.shape) # (N, 80513), a SciPy CSR sparse matrix
print(model.explain(activations[0], top_n=10))
model.explain(row, top_n=10) returns up to 10 active concepts from one
sparse output row, ordered by decreasing activation. Each result is a
dictionary with concept_id, label, and activation. It looks up the saved
labels; it does not run another model or make an API call. top_n controls
how many concepts are displayed, not the encoder's top-128 selection.
Use device="cuda" for an NVIDIA GPU. Keep the model in float32.
For large datasets, model.encode_iter(embeddings, batch_size=64) yields
sparse batches without collecting the full output in memory.
Example
With model loaded above, embed a sentence and display its 50 highest-activating
concepts. The embedding step uses the
Google GenAI SDK and requires your
Gemini API key in the GEMINI_API_KEY environment variable:
pip install google-genai
import os
from google import genai
from google.genai import types
text = "Djokovic Overwhelms Nadal for the Wimbledon Title"
client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
response = client.models.embed_content(
model="gemini-embedding-2-preview",
contents=[text],
config=types.EmbedContentConfig(
task_type="RETRIEVAL_DOCUMENT",
output_dimensionality=3072,
),
)
example_embeddings = np.array(
[embedding.values for embedding in response.embeddings], dtype=np.float32
)
example_activations = model.encode(example_embeddings)
for concept in model.explain(example_activations[0], top_n=50):
print(
f"{concept['concept_id']:5d} "
f"{concept['activation']:.6f} {concept['label']}"
)
The output lists the concept ID, activation score, and saved label for each of the top 50 active concepts, ordered from highest to lowest activation.
Concepts and prevalences
labels.csvhas one row per output coordinate:concept_idis the column index (0–80512),labelis the concept name, anddiscussionis a description of activating examples used to produce the label.prevalences.csvandprevalences.npzcontain the same heuristic corpus prevalences at activation thresholds>0.05and>0.1, with counts and one-sided 99% upper confidence bounds. These can be used to identify concepts that are enriched in your dataset.
- Downloads last month
- 68