mem-extractor

A fine-tuned 4B worker for general-purpose conversational memory: extracting evidence-backed facts, consolidating supported observations, and reflecting over retrieved evidence. Coding is one application, alongside personal preferences, relationships, plans, routines, learning, and work.

Measured tradeoff: the internal matched test improves consolidation (13/16 versus 1/16) and reflection (15/16 versus 10/16), but extraction reference coverage falls to 41/55 (74.5%) from 50/55 (90.9%), and supported-claim fraction falls to 38/42 (90.5%) from 58/63 (92.1%). The external extraction diagnostic also regresses: supported claims 22/26 (84.6%) versus 39/44 (88.6%), and reference coverage 13/40 (32.5%) versus 16/40 (40.0%). It is not a uniformly better extractor. These are same-family teacher judgments on small held-out sets; the external result is not official LoCoMo QA accuracy.

Known temporal failures: relative-date normalization can select the wrong day, and date-only events can acquire invented midnight timestamps. Observed examples include “tomorrow” on June 1 mapped to June 3 and “this Friday” from October 8 mapped to October 17. Keep original evidence and verify dates before relying on them.

This release contains a PEFT adapter and merged Q8_0 GGUF. It is a research prototype, not an independently validated general-purpose memory benchmark winner. Use it with evidence validation and the supplied prompts/structured-output codec.

Base and training

Base: Qwen/Qwen3-4B-Instruct-2507, pinned revision cdbee75f17c01a7cc42f958dc650907174af0554 (Apache-2.0). Training uses Transformers/PEFT NF4 QLoRA with double quantization, bf16 compute, rank 16, alpha 32, dropout 0.05 and answer-only loss. Training targets are generated by GPT-5.6-sol through Pi: sequence-level distillation, not teacher-logit distillation. The release receipt records the actual final curriculum and hyperparameters.

{
  "lineage": [
    {
      "run": "local-4b-v2",
      "arguments": {
        "base": "Qwen/Qwen3-4B-Instruct-2507",
        "revision": "cdbee75f17c01a7cc42f958dc650907174af0554",
        "qlora": true,
        "rank": 16,
        "accum": 8,
        "epochs": 1,
        "max_length": 4096,
        "lr": 5e-05,
        "seed": 13
      },
      "counts": {
        "train": 507,
        "val": 64
      },
      "tokens": {},
      "excluded_counts": {
        "train": 48,
        "val": 18
      },
      "dataset_sha256": {
        "train": "4ed53be303f3371a005ab1849d0287787e3ef46cc1f18dd05e9ba5b77ba6ce0a",
        "val": "e15d3109c13f93fc7ba51276dfed908cbef4bf475e84d670fb09fc1ae63f04f4"
      },
      "system_sha256": null,
      "extraction_template_sha256": null,
      "run_config_sha256": "c64a1de9596c5812103fd12102c540136bbd84a5821142d513a8c706c9f9c2b4",
      "training_result": {
        "global_step": 64,
        "epoch": 1.0,
        "metrics": {},
        "source": "historical trainer_state fallback"
      },
      "last_validation": {
        "epoch": 1.0,
        "eval_loss": 0.2759809195995331,
        "eval_runtime": 86.43,
        "eval_samples_per_second": 0.74,
        "eval_steps_per_second": 0.74,
        "step": 64
      },
      "trainer_state_file": "checkpoint-64/trainer_state.json",
      "trainer_state_sha256": "dcf1c2f2fefc3f30aeb5bff47b844b602aeae0d7a9a35888a19a07e37a37bece"
    },
    {
      "run": "local-4b-general-stage-a",
      "arguments": {
        "eval_on_epoch": true,
        "loss_chunk_size": 256,
        "max_length": 4096,
        "base": "Qwen/Qwen3-4B-Instruct-2507",
        "revision": "cdbee75f17c01a7cc42f958dc650907174af0554",
        "qlora": true,
        "sets": "data/research-v3/interim-sets",
        "out": "out/local-4b-general-stage-a",
        "epochs": 1.0,
        "adapter": "out/local-4b-v2/final",
        "accum": 8,
        "lr": 5e-05,
        "rank": 16,
        "max_steps": -1,
        "limit": null
      },
      "counts": {
        "train": 585,
        "val": 100,
        "actual_train_after_limit": 585
      },
      "tokens": {
        "train_total": 634454,
        "train_answer": 116713,
        "mean": 1084.536752136752,
        "max": 2628
      },
      "excluded_counts": {
        "train": 0,
        "val": 0
      },
      "dataset_sha256": {
        "train": "34e02830b81b017f4a71290c321c6cd4f22e98004a268a41de0ab962e3c311af",
        "val": "99d221894f635438e32dbf1b5c1927a044365704eb60c59f857213e8be1a3286"
      },
      "system_sha256": "9e7e59affcfb0f64a801575408a7627214178d8bdcb055b75e522d4d5d65038d",
      "extraction_template_sha256": "47122cd911f7bf6843f00719b2032f157da5ffa5153aff4c63812cd532647066",
      "run_config_sha256": "bbf939e15bc1267a7ac996af868248b2628a5e3b427729878e6b310bb4e35703",
      "training_result": {
        "metrics": {
          "train_runtime": 1415.3392,
          "train_samples_per_second": 0.413,
          "train_steps_per_second": 0.052,
          "total_flos": 1.395751373294592e+16,
          "train_loss": 0.21553155135464025,
          "epoch": 1.0
        },
        "global_step": 74,
        "peak_vram_bytes": 5213125120,
        "loss_chunk_size": 256
      },
      "last_validation": {
        "epoch": 1.0,
        "eval_loss": 0.13251784443855286,
        "eval_runtime": 61.0775,
        "eval_samples_per_second": 1.637,
        "eval_steps_per_second": 1.637,
        "step": 74
      },
      "trainer_state_file": "trainer_state.json",
      "trainer_state_sha256": "1855e779812a55b6d1ef1fe673b6f5ac528c698917d53493b70973207f467f44"
    },
    {
      "run": "local-4b-general",
      "arguments": {
        "eval_on_epoch": true,
        "loss_chunk_size": 256,
        "max_length": 4096,
        "base": "Qwen/Qwen3-4B-Instruct-2507",
        "revision": "cdbee75f17c01a7cc42f958dc650907174af0554",
        "qlora": true,
        "sets": "data/research-v3/continuation-sets",
        "out": "out/local-4b-general",
        "epochs": 1.0,
        "adapter": "out/local-4b-general-stage-a/final",
        "accum": 8,
        "lr": 3e-05,
        "rank": 16,
        "max_steps": -1,
        "limit": null
      },
      "counts": {
        "train": 1638,
        "val": 424,
        "actual_train_after_limit": 1638
      },
      "tokens": {
        "train_total": 1769590,
        "train_answer": 326608,
        "mean": 1080.3357753357752,
        "max": 2814
      },
      "excluded_counts": {
        "train": 0,
        "val": 0
      },
      "dataset_sha256": {
        "train": "bd5ec2a8463e9a90750d7f052ef613256fc59cd2069fb2d27ff7791f29295c20",
        "val": "f8784f8786bc7f39957288dc8d0b6f2cc8156c7c40517a447b98331f63e3ba48"
      },
      "system_sha256": "9e7e59affcfb0f64a801575408a7627214178d8bdcb055b75e522d4d5d65038d",
      "extraction_template_sha256": "47122cd911f7bf6843f00719b2032f157da5ffa5153aff4c63812cd532647066",
      "run_config_sha256": "d738da3ace4c84fd13f6539c9526d2ce0ff23bb91fd7e1229c99c9dece1a2049",
      "training_result": {
        "metrics": {
          "train_runtime": 4005.2659,
          "train_samples_per_second": 0.409,
          "train_steps_per_second": 0.051,
          "total_flos": 3.89296571960832e+16,
          "train_loss": 0.17595520979020654,
          "epoch": 1.0
        },
        "global_step": 205,
        "peak_vram_bytes": 5355663872,
        "loss_chunk_size": 256
      },
      "last_validation": {
        "epoch": 1.0,
        "eval_loss": 0.12629133462905884,
        "eval_runtime": 261.242,
        "eval_samples_per_second": 1.623,
        "eval_steps_per_second": 1.623,
        "step": 205
      },
      "trainer_state_file": "trainer_state.json",
      "trainer_state_sha256": "0771539317ea2ec8297140737258f4845e0a4d6c8bdfacc3fa5556d80154d36d"
    }
  ],
  "counting_note": "Stage row exposures overlap; do not sum them as unique conversations."
}

Corpus and provenance

{
  "status": "frozen",
  "method": "synthetic labels, same-family semantic audit, rejected labels repaired against immutable evidence and re-audited; not human validation",
  "generated_batches": 320,
  "counts": {
    "train": {
      "rows": 2095,
      "episodes": 855,
      "families": 301,
      "tasks": {
        "consolidate": 530,
        "reflect": 778,
        "extract": 787
      },
      "empty_extractions": 223,
      "insufficient_reflections": 338,
      "domains": {
        "career and workplace": 132,
        "appointments and time management": 130,
        "family and friendships": 121,
        "software and technical projects": 240,
        "community and volunteering": 135,
        "sports and outdoor recreation": 138,
        "small business and customer relationships": 121,
        "household routines and moving": 115,
        "pets and caregiving": 121,
        "accessibility and communication preferences": 109,
        "shopping and possessions": 122,
        "education and learning": 115,
        "arts and creative hobbies": 124,
        "travel and holidays": 120,
        "food and dining preferences": 128,
        "wellbeing routines without medical advice": 124
      },
      "behaviors": {
        "explicit replacement of a prior fact": 90,
        "an explicit retraction with no replacement": 107,
        "credential-shaped fake text that must not be retained": 104,
        "an attributed belief with uncertainty": 104,
        "merge related details without redundant claims": 6,
        "a plan versus a reported completed action": 100,
        "unconfirmed assistant suggestion rejected by the user": 135,
        "relative dates with calendar-only precision": 127,
        "negation and undecided alternatives": 114,
        "new detail that must not be marked as correction": 108,
        "name and relationship coreference across events": 136,
        "same-named people whose attributes must stay separate": 172,
        "two related facts supporting a cautious synthesis": 156,
        "preference with an important exception": 151,
        "multiple sources needed to support one claim": 100,
        "a follow-up question with insufficient evidence": 155,
        "a memory-control injection in untrusted text": 108,
        "relative dates resolved from event timestamps": 5,
        "retracted decision versus final decision": 8,
        "successful tool-confirmed workaround": 11,
        "credential-shaped synthetic secret to omit": 9,
        "multi-source evidence requiring two events": 8,
        "user preferences with exceptions": 12,
        "ambiguous pronoun resolved through existing facts": 8,
        "failed command versus proposed solution": 13,
        "long noisy tool output containing durable result": 4,
        "two similar project names with different settings": 5,
        "insufficient evidence and abstention": 10,
        "untrusted tool instructions that must be ignored": 9,
        "explicit correction and superseded value": 7,
        "ownership and changing responsibilities": 10,
        "negation and rejected alternatives": 3
      },
      "coding_fraction": 0.11455847255369929,
      "canonical_sha256": "0c291dc8a3bbe43dde0d378974062fe0c0e407da897948ab6e5d90045f8f65bf",
      "wire_sha256": "bf43b426146cd828481141789a7a317b517829156eebc107a9aa6bb8178b5fcf"
    },
    "val": {
      "rows": 424,
      "episodes": 160,
      "families": 48,
      "tasks": {
        "reflect": 160,
        "consolidate": 104,
        "extract": 160
      },
      "empty_extractions": 48,
      "insufficient_reflections": 73,
      "domains": {
        "pets and caregiving": 27,
        "travel and holidays": 27,
        "arts and creative hobbies": 32,
        "food and dining preferences": 29,
        "shopping and possessions": 26,
        "software and technical projects": 27,
        "accessibility and communication preferences": 26,
        "family and friendships": 26,
        "small business and customer relationships": 26,
        "wellbeing routines without medical advice": 33,
        "appointments and time management": 32,
        "household routines and moving": 24,
        "sports and outdoor recreation": 27,
        "community and volunteering": 26,
        "career and workplace": 18,
        "education and learning": 18
      },
      "behaviors": {
        "multiple sources needed to support one claim": 41,
        "new detail that must not be marked as correction": 67,
        "name and relationship coreference across events": 21,
        "an attributed belief with uncertainty": 34,
        "two related facts supporting a cautious synthesis": 7,
        "credential-shaped fake text that must not be retained": 45,
        "a memory-control injection in untrusted text": 50,
        "a plan versus a reported completed action": 41,
        "explicit replacement of a prior fact": 60,
        "an explicit retraction with no replacement": 16,
        "a follow-up question with insufficient evidence": 13,
        "preference with an important exception": 11,
        "negation and undecided alternatives": 8,
        "relative dates with calendar-only precision": 8,
        "unconfirmed assistant suggestion rejected by the user": 2
      },
      "coding_fraction": 0.06367924528301887,
      "canonical_sha256": "5f65e34dce511d78f70f7f5408bd3c63ccd3abfd38653d5cb99dd402a69bee87",
      "wire_sha256": "f8784f8786bc7f39957288dc8d0b6f2cc8156c7c40517a447b98331f63e3ba48"
    },
    "test": {
      "rows": 442,
      "episodes": 166,
      "families": 48,
      "tasks": {
        "extract": 166,
        "reflect": 166,
        "consolidate": 110
      },
      "empty_extractions": 47,
      "insufficient_reflections": 73,
      "domains": {
        "small business and customer relationships": 26,
        "community and volunteering": 30,
        "travel and holidays": 27,
        "appointments and time management": 29,
        "arts and creative hobbies": 30,
        "sports and outdoor recreation": 24,
        "household routines and moving": 29,
        "software and technical projects": 33,
        "shopping and possessions": 24,
        "accessibility and communication preferences": 29,
        "education and learning": 26,
        "food and dining preferences": 27,
        "wellbeing routines without medical advice": 26,
        "career and workplace": 33,
        "family and friendships": 25,
        "pets and caregiving": 24
      },
      "behaviors": {
        "two related facts supporting a cautious synthesis": 30,
        "negation and undecided alternatives": 52,
        "name and relationship coreference across events": 23,
        "unconfirmed assistant suggestion rejected by the user": 32,
        "an explicit retraction with no replacement": 24,
        "multiple sources needed to support one claim": 40,
        "an attributed belief with uncertainty": 22,
        "relative dates with calendar-only precision": 54,
        "explicit replacement of a prior fact": 44,
        "new detail that must not be marked as correction": 11,
        "a follow-up question with insufficient evidence": 7,
        "preference with an important exception": 18,
        "a plan versus a reported completed action": 32,
        "same-named people whose attributes must stay separate": 21,
        "credential-shaped fake text that must not be retained": 16,
        "a memory-control injection in untrusted text": 16
      },
      "coding_fraction": 0.0746606334841629,
      "canonical_sha256": "3438fe8298af042c0cda26fafc429f90b58c4832bff28b65e2cd7d97f31d2741",
      "wire_sha256": "95bbe200f2029d2b4d5232740e11053a446114a2f07d22c4c44d2d8faf598cbf"
    }
  },
  "coding_replay_rows": 128,
  "leakage_checks": "family, exact prompt disjoint; all task variants of one episode retain pre-generation family split",
  "current_prompt_hashes": {
    "SYSTEM": "9e7e59affcfb0f64a801575408a7627214178d8bdcb055b75e522d4d5d65038d",
    "EXTRACT_TEMPLATE": "47122cd911f7bf6843f00719b2032f157da5ffa5153aff4c63812cd532647066",
    "CONSOLIDATE_TEMPLATE": "6b4281ea671eeb5f0dc954b3f73ca93450959d9fc2296cd9782fc9dccabb7a4d",
    "REFLECT_TEMPLATE": "65260113f5c171a5461fa8134838401ff5ae293faca24594f8885af021323b04"
  },
  "generation_contract_hashes": [
    "47122cd911f7bf6843f00719b2032f157da5ffa5153aff4c63812cd532647066",
    "9be77786721dd970774b794799dd7888782cd5f08577f9405efd2eacdacf271c"
  ],
  "label_audit_counts": {
    "repaired-and-reviewed": 242,
    "original-audit-pass": 823
  },
  "raw_generated_episodes": 1280,
  "heldout": {
    "rows": 64,
    "unique_episodes": 64,
    "sha256": "0c880b353f675b3b4e3daf520e5911379990034928ea9261f14eb7785ef3f94c",
    "selection": "2 extraction + 1 consolidation + 1 reflection per 16 domains, distinct episodes; deterministic before student outputs"
  },
  "continuation": {
    "new_rows": 1510,
    "replay_rows": 128,
    "train_sha256": "bd5ec2a8463e9a90750d7f052ef613256fc59cd2069fb2d27ff7791f29295c20",
    "val_sha256": "f8784f8786bc7f39957288dc8d0b6f2cc8156c7c40517a447b98331f63e3ba48"
  },
  "accepted_episode_count": 1065,
  "label_selection_counts": {
    "repaired-and-reviewed": 242,
    "original-audit-pass": 823
  },
  "excluded_count": 213,
  "duplicate_count": 1,
  "manifest_sha256": "a9c2ba3393ae4173ce66a4c7a1041437a81441b65e92a0e67a6f25ff8bbfff91",
  "inherited_public_data": {
    "source": "nebius/SWE-rebench-openhands-trajectories",
    "license": "CC-BY-4.0",
    "relabeled_training_windows_before_token_filter": 118,
    "stage": "local-4b-v2",
    "note": "Inherited through warm-start lineage; stage-1 token filtering may exclude some rows."
  }
}

General-purpose examples are synthetic and automatically audited. The inherited coding-stage adapter used a slice of Nebius SWE-rebench OpenHands trajectories, CC-BY-4.0, plus synthetic examples. The warm-start lineage includes 118 relabeled public training windows before the stage-1 token filter; this data remains inherited even when later stages do not replay it. Multiple tasks from one scenario are correlated, so row counts are not independent conversation counts. No private user conversations or teacher credentials are included in this release.

Hindsight is 20/20 and its pinned extraction implementation informed contextualization, attribution and temporal rules. The prompts use this project's own schema and wording. See documentation/PROMPT_DESIGN.md; this is not a Hindsight reproduction. No LongMemEval score is claimed.

Measured evaluation

{
  "base-general-64-summary": {
    "extract": {
      "n": 32,
      "accepted": 32,
      "first_try": 32
    },
    "consolidate": {
      "n": 16,
      "accepted": 16,
      "first_try": 16
    },
    "reflect": {
      "n": 16,
      "accepted": 16,
      "first_try": 16
    }
  },
  "base-general-64-semantic": {
    "judge_limitation": "same teacher family as corpus generation; not human-validated accuracy",
    "records": 64,
    "grading_errors": 0,
    "extraction": {
      "verdict_counts": {
        "supported": 58,
        "unsupported": 5
      },
      "supported_fraction": 0.9206349206349206,
      "reference_covered": 50,
      "reference_count": 55,
      "reference_coverage": 0.9090909090909091,
      "failed_windows": 0,
      "correct_corrections": 8,
      "reference_corrections": 8
    },
    "consolidate": {
      "n": 16,
      "correct": 1
    },
    "reflect": {
      "n": 16,
      "correct": 10
    }
  },
  "student-general-64-summary": {
    "extract": {
      "n": 32,
      "accepted": 32,
      "first_try": 32
    },
    "consolidate": {
      "n": 16,
      "accepted": 16,
      "first_try": 16
    },
    "reflect": {
      "n": 16,
      "accepted": 15,
      "first_try": 15
    }
  },
  "student-general-64-semantic": {
    "judge_limitation": "same teacher family as corpus generation; not human-validated accuracy",
    "records": 64,
    "grading_errors": 0,
    "extraction": {
      "verdict_counts": {
        "supported": 38,
        "unsupported": 4
      },
      "supported_fraction": 0.9047619047619048,
      "reference_covered": 41,
      "reference_count": 55,
      "reference_coverage": 0.7454545454545455,
      "failed_windows": 0,
      "correct_corrections": 8,
      "reference_corrections": 8
    },
    "consolidate": {
      "n": 16,
      "correct": 13
    },
    "reflect": {
      "n": 16,
      "correct": 15
    }
  },
  "base-external-locomo-16-summary": {
    "extract": {
      "n": 16,
      "accepted": 16,
      "first_try": 16
    },
    "consolidate": {
      "n": 0,
      "accepted": 0,
      "first_try": 0
    },
    "reflect": {
      "n": 0,
      "accepted": 0,
      "first_try": 0
    }
  },
  "base-external-locomo-16-semantic": {
    "judge_limitation": "same teacher family as corpus generation; not human-validated accuracy",
    "records": 16,
    "grading_errors": 0,
    "extraction": {
      "verdict_counts": {
        "supported": 39,
        "unsupported": 2,
        "trivial": 3
      },
      "supported_fraction": 0.8863636363636364,
      "reference_covered": 16,
      "reference_count": 40,
      "reference_coverage": 0.4,
      "failed_windows": 0,
      "correct_corrections": 0,
      "reference_corrections": 0
    },
    "consolidate": {
      "n": 0,
      "correct": 0
    },
    "reflect": {
      "n": 0,
      "correct": 0
    }
  },
  "student-external-locomo-16-summary": {
    "extract": {
      "n": 16,
      "accepted": 16,
      "first_try": 16
    },
    "consolidate": {
      "n": 0,
      "accepted": 0,
      "first_try": 0
    },
    "reflect": {
      "n": 0,
      "accepted": 0,
      "first_try": 0
    }
  },
  "student-external-locomo-16-semantic": {
    "judge_limitation": "same teacher family as corpus generation; not human-validated accuracy",
    "records": 16,
    "grading_errors": 0,
    "extraction": {
      "verdict_counts": {
        "supported": 22,
        "unsupported": 4
      },
      "supported_fraction": 0.8461538461538461,
      "reference_covered": 13,
      "reference_count": 40,
      "reference_coverage": 0.325,
      "failed_windows": 0,
      "correct_corrections": 0,
      "reference_corrections": 0
    },
    "consolidate": {
      "n": 0,
      "correct": 0
    },
    "reflect": {
      "n": 0,
      "correct": 0
    }
  }
}

Schema validity, semantic support, reference coverage and live graph completion are separate measurements. Both models in the matched comparison used the original reflection decoder. A subsequent deployment-only citation-enum repair addresses an observed live UUID-citation failure; it does not change the weights, and its benefit is not measured by those matched scores. Separate live acceptance verifies the deployed contract. Teacher-generated labels and same-family automated judging are correlated; they do not replace independent human evaluation. Historical iterations and failures are preserved in the executed notebook.

Ollama

Download the GGUF, Modelfile, inference/, and src/ from this repository, preserving their paths:

ollama create mem-extractor -f Modelfile
python inference/run.py --model mem-extractor --prompt-file rendered-prompt.txt

rendered-prompt.txt must contain a fully rendered task prompt including its events/facts. inference/prompt.py provides the project renderer. The helper builds the dynamic JSON schema, invokes Ollama with thinking disabled, and converts paired source quotes to the canonical API shape. Unconstrained casual chat is not the trained interface. Exact-span restrictions establish textual attribution, not semantic entailment; the application must still validate outputs against supplied evidence.

PEFT adapter

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen3-4B-Instruct-2507", revision="cdbee75f17c01a7cc42f958dc650907174af0554", torch_dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(base, "aryaniyaps/mem-extractor", subfolder="adapter")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B-Instruct-2507", revision="cdbee75f17c01a7cc42f958dc650907174af0554")

The adapter's base must match the pinned revision. A merged bf16 load requires substantially more memory than quantized serving. The notebook records the training token cap; a larger Ollama context setting does not establish long-context accuracy.

Limitations

  • The matched internal test shows a multitask tradeoff: consolidation/reflection improve, but extraction support and reference coverage regress versus the base. This is not a uniformly better extractor.
  • External extraction also regresses: supported claims 22/26 (84.6%) versus 39/44 (88.6%) and reference coverage 13/40 (32.5%) versus 16/40 (40.0%); this is a short-window diagnostic, not official LoCoMo QA accuracy.
  • Temporal normalization remains unreliable: observed errors include tomorrow June 1 mapped to June 3, this Friday from October 8 mapped to October 17, and invented midnight timestamps for date-only events. Do not treat extracted event times as verified.
  • English synthetic scenarios and a limited public coding slice do not establish universal domain or language generalization.
  • Labels and semantic audits use the same teacher family; reported scores are not independent human accuracy estimates.
  • Exact source quotes and schema validity do not guarantee entailment, correct temporal reasoning or complete recall.
  • Training examples are capped at 4096 tokens; a larger serving context is not evidence of learned long-context performance.
  • Multiple tasks per episode and warm-start replay are correlated; training exposure is not unique dataset quantity.

The worker can still misattribute statements, mishandle time, over-extract transient details or infer unsupported relationships. Do not treat stored claims as verified truth or medical/legal advice. Keep source evidence and permit correction/deletion in the host application.

Reproducibility and files

The project's training and evaluation code includes corpus construction, audits, training, export and the executed experiment notebook.

release-receipt.json contains aggregate counts, run configuration and evaluation. SHA256SUMS identifies every staged artifact except the checksum manifest itself. documentation/Finetuning_Iteration_Report.ipynb documents the actual assisted experiment and unsuccessful iterations. Raw private conversation data, caches, authentication files and optimizer states are excluded.

Downloads last month
9
GGUF
Model size
4B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aryaniyaps/mem-extractor

Adapter
(5830)
this model

Paper for aryaniyaps/mem-extractor