H&M recommendation tools
Frozen H&M catalog artifacts and behavior-trained recommendation tools.
The category and run layout follows iaouali/amazon-tools.
This repository contains tool artifacts, not a standalone Transformers model
or a recommendation-agent LLM.
Layout
H_and_M/
data/
product_texts.json
subcategories.json
embeddings/
item_embeddings_qwen.pt
manifest.json
runs/shared-train-best-recipes-v1/
models/recsys_model_qwen.pt
embeddings/comp_embeddings_qwen.pt
metrics/retrieval_baseline_qwen.json
metrics/retrieval_history_qwen.json
training_recipe.json
manifest.json
release_manifest.json
All five catalog/model/embedding files are byte-identical to the frozen inference artifacts used in the H&M experiments. Manifests give sizes and SHA-256 checksums. Preserve zero-padded article IDs as strings, and keep catalog and embedding order aligned.
Trained retrieval
The retrieval checkpoint is a query-side LoRA adapter for
Qwen/Qwen3-Embedding-0.6B, selected at epoch 5 by internal validation
multi-positive MRR. The selected score is 0.18585462868213654.
Training uses positive interactions from shared_train and an internal
product-disjoint retrieval validation split. It does not use hidden queries.
The checkpoint contains adapter_state_dict, model_config, epoch and
validation metadata, and the original optimizer/scheduler state. The base
embedding backbone and tokenizer are not bundled. LoRA has rank 32,
alpha 32, dropout 0.1, and targets attention q/k/v/o plus MLP gate/up/down
projections. Query length is 256 and item length is 512. Query instruction:
"Given a shopping query, retrieve relevant products."
from huggingface_hub import hf_hub_download
import torch
# For reproducibility, replace main with the immutable repository commit SHA.
path = hf_hub_download(
repo_id="iaouali/hm-tools",
filename="H_and_M/runs/shared-train-best-recipes-v1/models/recsys_model_qwen.pt",
revision="main",
)
checkpoint = torch.load(path, map_location="cpu", weights_only=True)
adapter_state = checkpoint["adapter_state_dict"]
model_config = checkpoint["model_config"]
This is a custom PyTorch checkpoint, not a PEFT adapter_model.safetensors
directory. Instantiate the matching query-encoder/LoRA architecture before
loading its adapter state. Use the same frozen item vectors to score queries.
Do not point AutoModel.from_pretrained at this repository.
Catalog and complementarity
The catalog contains 104,993 articles. Base item vectors have 1,024 dimensions
and are normalized. item_embeddings_qwen.pt contains item_ids and
item_embs. The trained complementarity file contains the same item_ids,
plus normalized anchor_embs and complement_embs, both 1,024-dimensional.
Catalog pair scores are dot products of the corresponding anchor and
complement vectors, with the tool's subcategory exclusions applied separately.
The complementarity vectors are the frozen residual-projection variant used
by the evaluated H&M agent (recorded epoch-2 selection), not the earlier
256-dimensional MLP variant. The separate residual projection-head checkpoint
was not recovered for this publication and is not included. These vectors
support scoring the existing catalog; they do not enable projecting new
products or resuming complementarity training. No different checkpoint is
substituted under comp_model_qwen.pt.
Semantic retrieval disables the query adapter. Semantic complementarity and near-duplicate pruning use the base item vectors. Near-duplicate pruning has no separate trained checkpoint. A compatible tool implementation is required for retrieval and response formatting; this repository supplies its artifacts.
Data and scope
Query/reference sets are separate: https://huggingface.co/datasets/iaouali/hm-benchmark. Raw transactions, user histories, user-split keys, raw user identifiers, agent weights, RL policies, and reward models are not included here. The run name follows Amazon's layout; it does not imply that H&M uses Amazon's complementarity architecture or that a complete retraining release is provided.
The underlying dataset originates from the H&M Personalized Fashion Recommendations competition: https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations. Source data and base models remain subject to their respective terms. No new blanket license is granted for third-party catalog content or model weights.