Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
bng
dosmot13
1
3
Follow
0 followers
Ā·
3 following
AI & ML interests
None yet
Recent Activity
upvoted
an
article
2 days ago
Rebuilding AUTOMATIC1111 with Gradio Workflow
reacted
to
mayafree
's
post
with š
2 days ago
JEV Ecosystems ā every answer-verification vendor publishes a benchmark, and every one of them wins it. So we ran 13 of them on one test set: 2,018 items, identical labels, same grading code. šÆ Leaderboard https://huggingface.co/spaces/mayafree/typed-decision-leaderboard š Full write-up (method, mechanism, limits) https://huggingface.co/blog/mayafree/jve-ecosystems š§Ŗ Try it ā ZTC, JEV and Laya on the same input, side by side https://huggingface.co/spaces/mayafree/verifier-playground Three results 1ļøā£ Only three systems clear 0.70 ā ZTC (397B) 0.7364 Ā· JEV 0.7350 Ā· ZTC (27B) 0.7282. First and second differ by 0.0014, so no rank is assigned. 2ļøā£ A baseline that reads nothing but answer length and formatting scores 0.7036. Eight of the thirteen fall below it. A leaderboard without that line is flattering its entrants. 3ļøā£ Bigger does not win. On scientific reasoning, 27B 0.7410 beats 397B 0.6287 ā a model fourteen times larger scoring 0.11 lower. And AUC is not the number you deploy on. Same 20% retry budget, wired into an agent loop, against a 74.83% no-gate baseline: ZTC +1.34 pp Ā· JEV ā0.07 pp Ā· random ā0.25 pp. The mechanism is the interesting part. Re-answering is double-edged: 38% of wrong answers get fixed, and 30% of right answers get broken. So a gate is paid for by precision, not recall. Of the 403 items JEV routed for a retry, 216 were already correct. 0.0014 AUC apart; 1.4 points of end-to-end agent accuracy apart. Scores, labels and grading code are published in full. Four public reproductions that would not run from their released artefacts are listed too, with the failure and a link, and no score. Don't take the table's word for it ā paste your own case into the playground and watch all three answer at once. Want a system added? Open a discussion on the Space.
liked
a Space
about 2 months ago
fancyfeast/llama-bigasp-prompt-enhancer
View all activity
Organizations
None yet
models
1
dosmot13/Kr3a
13B
ā¢
Updated
2 days ago
datasets
0
None public yet