H&M recommendation tools

Frozen H&M catalog artifacts and behavior-trained recommendation tools. The category and run layout follows iaouali/amazon-tools. This repository contains tool artifacts, not a standalone Transformers model or a recommendation-agent LLM.

Layout

H_and_M/
  data/
    product_texts.json
    subcategories.json
  embeddings/
    item_embeddings_qwen.pt
  manifest.json
  runs/shared-train-best-recipes-v1/
    models/recsys_model_qwen.pt
    embeddings/comp_embeddings_qwen.pt
    metrics/retrieval_baseline_qwen.json
    metrics/retrieval_history_qwen.json
    training_recipe.json
    manifest.json
    release_manifest.json

All five catalog/model/embedding files are byte-identical to the frozen inference artifacts used in the H&M experiments. Manifests give sizes and SHA-256 checksums. Preserve zero-padded article IDs as strings, and keep catalog and embedding order aligned.

Trained retrieval

The retrieval checkpoint is a query-side LoRA adapter for Qwen/Qwen3-Embedding-0.6B, selected at epoch 5 by internal validation multi-positive MRR. The selected score is 0.18585462868213654. Training uses positive interactions from shared_train and an internal product-disjoint retrieval validation split. It does not use hidden queries.

The checkpoint contains adapter_state_dict, model_config, epoch and validation metadata, and the original optimizer/scheduler state. The base embedding backbone and tokenizer are not bundled. LoRA has rank 32, alpha 32, dropout 0.1, and targets attention q/k/v/o plus MLP gate/up/down projections. Query length is 256 and item length is 512. Query instruction: "Given a shopping query, retrieve relevant products."

from huggingface_hub import hf_hub_download
import torch

# For reproducibility, replace main with the immutable repository commit SHA.
path = hf_hub_download(
    repo_id="iaouali/hm-tools",
    filename="H_and_M/runs/shared-train-best-recipes-v1/models/recsys_model_qwen.pt",
    revision="main",
)
checkpoint = torch.load(path, map_location="cpu", weights_only=True)
adapter_state = checkpoint["adapter_state_dict"]
model_config = checkpoint["model_config"]

This is a custom PyTorch checkpoint, not a PEFT adapter_model.safetensors directory. Instantiate the matching query-encoder/LoRA architecture before loading its adapter state. Use the same frozen item vectors to score queries. Do not point AutoModel.from_pretrained at this repository.

Catalog and complementarity

The catalog contains 104,993 articles. Base item vectors have 1,024 dimensions and are normalized. item_embeddings_qwen.pt contains item_ids and item_embs. The trained complementarity file contains the same item_ids, plus normalized anchor_embs and complement_embs, both 1,024-dimensional. Catalog pair scores are dot products of the corresponding anchor and complement vectors, with the tool's subcategory exclusions applied separately.

The complementarity vectors are the frozen residual-projection variant used by the evaluated H&M agent (recorded epoch-2 selection), not the earlier 256-dimensional MLP variant. The separate residual projection-head checkpoint was not recovered for this publication and is not included. These vectors support scoring the existing catalog; they do not enable projecting new products or resuming complementarity training. No different checkpoint is substituted under comp_model_qwen.pt.

Semantic retrieval disables the query adapter. Semantic complementarity and near-duplicate pruning use the base item vectors. Near-duplicate pruning has no separate trained checkpoint. A compatible tool implementation is required for retrieval and response formatting; this repository supplies its artifacts.

Data and scope

Query/reference sets are separate: https://huggingface.co/datasets/iaouali/hm-benchmark. Raw transactions, user histories, user-split keys, raw user identifiers, agent weights, RL policies, and reward models are not included here. The run name follows Amazon's layout; it does not imply that H&M uses Amazon's complementarity architecture or that a complete retraining release is provided.

The underlying dataset originates from the H&M Personalized Fashion Recommendations competition: https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations. Source data and base models remain subject to their respective terms. No new blanket license is granted for third-party catalog content or model weights.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for iaouali/hm-tools

Finetuned
(555)
this model