Modified Gemma weights. These are fine-tunes of google/codegemma-2b, not the original CodeGemma weights. Gemma is provided under and subject to the Gemma Terms of Use; use is also subject to the Gemma Prohibited Use Policy. Those terms carry over to these checkpoints and to anything you distribute from them (see NOTICE).
Repl-Ghidra
Repl-Ghidra: CodeGemma-2B retrained on GenNm's gennm-ghidra-O0 corpus with the SFT + SymPO/DPO recipe (paper Section 4.2). Scores 41.89 / 39.78 per-binary precision/recall on GenNm's Ghidra-O0 test set (701 binaries); see RECIPE.txt for the exact training schedule.
Part of the ACSAC 2026 artifact for R+R: Revisiting LLM-Based Binary Name Recovery for Real-World Malware Analysis.
Layout: a single flat checkpoint (config.json, safetensors, tokenizer) -- no architecture subfolders. Input/output format follows GenNm: a Ghidra-decompiled function body
followed by Q:[FUN_...], answered with A:{'FUN_...': 'name', ...}. See the artifact README for the
inference driver (artifact/code/inference/eval_test.py) and scoring.
- Downloads last month
- -
Model tree for nghi85/Repl-Ghidra
Base model
google/codegemma-2b