wikimedia (Wikimedia)

posted an update about 1 month ago

Post

1361

🚨 How green is your model? 🌱 Introducing a new feature in the Comparator tool: Environmental Impact for responsible #LLM research!
👉 open-llm-leaderboard/comparator
Now, you can not only compare models by performance, but also by their environmental footprint!

🌍 The Comparator calculates CO₂ emissions during evaluation and shows key model characteristics: evaluation score, number of parameters, architecture, precision, type... 🛠️
Make informed decisions about your model's impact on the planet and join the movement towards greener AI!

resquito-wmf

in wikimedia/structured-wikipedia about 1 month ago

Reverts timestamp data type to string

1

#10 opened 3 months ago by

albertvillanova

resquito-wmf

updated a dataset about 1 month ago

wikimedia/structured-wikipedia

Preview • Updated Nov 14 • 836 • 54

resquito-wmf

in wikimedia/structured-wikipedia about 1 month ago

Upload README.md

#14 opened about 1 month ago by

resquito-wmf

albertvillanova

posted an update about 2 months ago

Post

1454

🚀 New feature of the Comparator of the 🤗 Open LLM Leaderboard: now compare models with their base versions & derivatives (finetunes, adapters, etc.). Perfect for tracking how adjustments affect performance & seeing innovations in action. Dive deeper into the leaderboard!

🛠️ Here's how to use it:
1. Select your model from the leaderboard.
2. Load its model tree.
3. Choose any base & derived models (adapters, finetunes, merges, quantizations) for comparison.
4. Press Load.
See side-by-side performance metrics instantly!

Ready to dive in? 🏆 Try the 🤗 Open LLM Leaderboard Comparator now! See how models stack up against their base versions and derivatives to understand fine-tuning and other adjustments. Easier model analysis for better insights! Check it out here: open-llm-leaderboard/comparator 🌐

albertvillanova

posted an update about 2 months ago

Post

3114

🚀 Exciting update! You can now compare multiple models side-by-side with the Hugging Face Open LLM Comparator! 📊

open-llm-leaderboard/comparator

Dive into multi-model evaluations, pinpoint the best model for your needs, and explore insights across top open LLMs all in one place. Ready to level up your model comparison game?

albertvillanova

posted an update about 2 months ago

Post

1220

🚨 Instruct-tuning impacts models differently across families! Qwen2.5-72B-Instruct excels on IFEval but struggles with MATH-Hard, while Llama-3.1-70B-Instruct avoids MATH performance loss! Why? Can they follow the format in examples? 📊 Compare models: open-llm-leaderboard/comparator

albertvillanova

posted an update 2 months ago

Post

1910

Finding the Best SmolLM for Your Project

Need an LLM assistant but unsure which hashtag#smolLM to run locally? With so many models available, how can you decide which one suits your needs best? 🤔

If the model you’re interested in is evaluated on the Hugging Face Open LLM Leaderboard, there’s an easy way to compare them: use the model Comparator tool: open-llm-leaderboard/comparator
Let’s walk through an example👇

Let’s compare two solid options:
- Qwen2.5-1.5B-Instruct from Alibaba Cloud Qwen (1.5B params)
- gemma-2-2b-it from Google (2.5B params)

For an assistant, you want a model that’s great at instruction following. So, how do these two models stack up on the IFEval task?

What about other evaluations?
Both models are close in performance on many other tasks, showing minimal differences. Surprisingly, the 1.5B Qwen model performs just as well as the 2.5B Gemma in many areas, even though it's smaller in size! 📊

This is a great example of how parameter size isn’t everything. With efficient design and training, a smaller model like Qwen2.5-1.5B can match or even surpass larger models in certain tasks.

Looking for other comparisons? Drop your model suggestions below! 👇

frimelle

authored a paper 2 months ago

Wikimedia data for AI: a review of Wikimedia datasets for NLP tasks and AI-assisted editing

Paper • 2410.08918 • Published Oct 11 • 2

albertvillanova

posted an update 2 months ago

Post

1946

🚨 We’ve just released a new tool to compare the performance of models in the 🤗 Open LLM Leaderboard: the Comparator 🎉
open-llm-leaderboard/comparator

Want to see how two different versions of LLaMA stack up? Let’s walk through a step-by-step comparison of LLaMA-3.1 and LLaMA-3.2. 🦙🧵👇

1/ Load the Models' Results
- Go to the 🤗 Open LLM Leaderboard Comparator: open-llm-leaderboard/comparator
- Search for "LLaMA-3.1" and "LLaMA-3.2" in the model dropdowns.
- Press the Load button. Ready to dive into the results!

2/ Compare Metric Results in the Results Tab 📊
- Head over to the Results tab.
- Here, you’ll see the performance metrics for each model, beautifully color-coded using a gradient to highlight performance differences: greener is better! 🌟
- Want to focus on a specific task? Use the Task filter to hone in on comparisons for tasks like BBH or MMLU-Pro.

3/ Check Config Alignment in the Configs Tab ⚙️
- To ensure you’re comparing apples to apples, head to the Configs tab.
- Review both models’ evaluation configurations, such as metrics, datasets, prompts, few-shot configs...
- If something looks off, it’s good to know before drawing conclusions! ✅

4/ Compare Predictions by Sample in the Details Tab 🔍
- Curious about how each model responds to specific inputs? The Details tab is your go-to!
- Select a Task (e.g., MuSR) and then a Subtask (e.g., Murder Mystery) and then press the Load Details button.
- Check out the side-by-side predictions and dive into the nuances of each model’s outputs.

5/ With this tool, it’s never been easier to explore how small changes between model versions affect performance on a wide range of tasks. Whether you’re a researcher or enthusiast, you can instantly visualize improvements and dive into detailed comparisons.

🚀 Try the 🤗 Open LLM Leaderboard Comparator now and take your model evaluations to the next level!

KingNish

posted an update 3 months ago

Post

6427

Realtime Whisper Large v3 Turbo Demo:
It transcribes audio in about 0.3 seconds.

KingNish/Realtime-whisper-large-v3-turbo

KingNish

posted an update 3 months ago

Post

6904

Exciting news! Introducing super-fast AI video assistant, currently in beta. With a minimum latency of under 500ms and an average latency of just 600ms.

DEMO LINK:
KingNish/Live-Video-Chat

1 reply

·

sdelbecque

updated a dataset 3 months ago

wikimedia/structured-wikipedia

Preview • Updated Nov 14 • 836 • 54

KingNish

posted an update 3 months ago

Post

3140

A super good and fast image inpainting demo is here.
Its' super cool and realistic.

Demo by @OzzyGT (Must try):
OzzyGT/diffusers-fast-inpaint

albertvillanova

posted an update 3 months ago

Post

1524

Check out the new Structured #Wikipedia dataset by Wikimedia Enterprise: abstract, infobox, structured sections, main image,...

Currently in early beta (English & French). Explore it and give feedback: wikimedia/structured-wikipedia

More info: https://enterprise.wikimedia.com/blog/hugging-face-dataset/
@sdelbecque @resquito-wmf

KingNish

posted an update 3 months ago

Post

3573

Mistral Nemo is better than many models in 1st grader level reasoning.

KingNish

posted an update 3 months ago

Post

3898

I am experimenting with Flux and trying to push it to its limits without training (as I am GPU-poor 😅).
I found some flaws in the pipelines, which I resolved, and now I am able to generate an approx similar quality image as Flux Schnell 4 steps in just 1 step.
Demo Link:
KingNish/Realtime-FLUX

1 reply

·

KingNish

posted an update 3 months ago

Post

1886

I am excited to announce a major speed updated in Voicee, a superfast voice assistant.

It has now achieved latency <250 ms.
While its average latency is about 500ms.
KingNish/Voicee

This become Possible due to newly launched @sambanovasystems cloud.

You can also use your own API Key to get fastest speed.
You can get on from here: https://cloud.sambanova.ai/apis

For optimal performance use Google Chrome.

Please try Voicee and share your valuable feedback to help me further improve its performance and usability.
Thank you!

KingNish

posted an update 4 months ago

Post

3588

Introducing Voicee, A superfast voice fast assistant.
KingNish/Voicee
It achieved latency <500 ms.
While its average latency is 700ms.
It works best in Google Chrome.
Please try and give your feedbacks.
Thank you. 🤗

3 replies

·

KingNish

posted an update 5 months ago

Post

5871

Introducing OpenCHAT mini: a lightweight, fast, and unlimited version of OpenGPT 4o.

KingNish/OpenCHAT-mini2

It has unlimited web search, vision and image generation.

Please take a look and share your review. Thank you! 🤗

7 replies

·

Wikimedia

AI & ML interests

Recent Activity

wikimedia's activity

Reverts timestamp data type to string

wikimedia/structured-wikipedia

Upload README.md

Wikimedia data for AI: a review of Wikimedia datasets for NLP tasks and AI-assisted editing

wikimedia/structured-wikipedia

AI & ML interests

Recent Activity

Team members 25

wikimedia's activity

Reverts timestamp data type to string

Upload README.md