Michael Leonardo Aguas

sensei-ml

AI & ML interests

Computer Vision Natural Language Processing Large Language Models

Recent Activity

updated a model 4 days ago

sensei-ml/llama_3.2_fine_tuned

published a model about 1 month ago

sensei-ml/llama_3.2_fine_tuned

reacted to louisbrulenaudet's post with ❤️ 7 months ago

My biggest release of the year: a series of 7 specialized embedding models for information retrieval within tax documents, is now available for free on Hugging Face 🤗 These new models aim to offer an open source alternative for in-domain semantic search from large text corpora and will improve RAG systems and context addition for large language models. Trained on more than 43 million tax tokens derived from semi-synthetic and raw-synthetic data, enriched by various methods (in particular MSFT's evol-instruct by @intfloat), and corrected by humans, this project is the fruit of hundreds of hours of work and is the culmination of a global effort to open up legal technologies that has only just begun. A big thank you to Microsoft for Startups for giving me access to state-of-the-art infrastructure to train these models, and to @julien-c, @clem 🤗, @thomwolf and the whole HF team for the inference endpoint API and the generous provision of Meta LLama-3.1-70B. Special thanks also to @tomaarsen for his invaluable advice on training embedding models and Loss functions ❤️ Models are available on my personal HF page, into the Lemone-embed collection: https://huggingface.co/collections/louisbrulenaudet/lemone-embed-66fdc24000df732b395df29b

View all activity

Organizations

sensei-ml's activity

updated a model 4 days ago

sensei-ml/llama_3.2_fine_tuned

Updated 4 days ago • 38

published a model about 1 month ago

sensei-ml/llama_3.2_fine_tuned

Updated 4 days ago • 38

reacted to louisbrulenaudet's post with ❤️ 7 months ago

Post

2155

My biggest release of the year: a series of 7 specialized embedding models for information retrieval within tax documents, is now available for free on Hugging Face 🤗

These new models aim to offer an open source alternative for in-domain semantic search from large text corpora and will improve RAG systems and context addition for large language models.

Trained on more than 43 million tax tokens derived from semi-synthetic and raw-synthetic data, enriched by various methods (in particular MSFT's evol-instruct by @intfloat ), and corrected by humans, this project is the fruit of hundreds of hours of work and is the culmination of a global effort to open up legal technologies that has only just begun.

A big thank you to Microsoft for Startups for giving me access to state-of-the-art infrastructure to train these models, and to @julien-c , @clem 🤗, @thomwolf and the whole HF team for the inference endpoint API and the generous provision of Meta LLama-3.1-70B. Special thanks also to @tomaarsen for his invaluable advice on training embedding models and Loss functions ❤️

Models are available on my personal HF page, into the Lemone-embed collection: louisbrulenaudet/lemone-embed-66fdc24000df732b395df29b