Daniel Vila's picture

Daniel Vila

dvilasuero

AI & ML interests

RLHF, RLAIF, DPO, data, data, data

Recent Activity

reacted to davanstrien's post with 🚀 1 day ago
The https://huggingface.co/datasets/data-is-better-together/fineweb-c dataset is growing! This week a few more languages have got 1,000 annotations for the educational quality of data from https://huggingface.co/datasets/HuggingFaceFW/fineweb-2. Why should you care? The quality of pre-training data can have a big impact on the performance of downstream language models trained on that data (https://huggingface.co/spaces/HuggingFaceFW/blogpost-fineweb-v1). Being able to filter by educational quality is on way of improving the quality of the data you use for training an LLM. Very importantly this approach can also reduce the amount of data needed for pertaining. Why not use an LLM? LLMs can be used to annotate educational quality for a subset of data. This data can then be used to train a smaller encoder only model to label the full dataset. However, this may not work well for languages outside of english. This is where fineweb-c (community) comes in. The community is annotating the educational quality of fineweb2 data. Currently 114 languages have some annotations. These annotations will enable a number of things: - Evaluate whether an LLM can label the educational quality for texts in that language well - Directly be used for training quality classifiers - Help discover other rules and huerisitcs for refining fineweb2 further for different languages. This week the following languages where done: Swedish thanks to: @Lauler @AntonVic @ohallstrom @bjarlestam @menbom @Ekgren @apsod Ukrainian thanks to: @hannayukhymenko @robinhad @realPivo @RabotiahovDmytro @reciprocate Assamese thanks to: @moyoor97 @Arpanjyoti @nawaf-helmi123 @pahigogoi1 @aelhence @kishorekashyap Want to learn more: https://huggingface.co/blog/davanstrien/fineweb2-community Contribute yourself here: https://huggingface.co/spaces/data-is-better-together/fineweb-c
liked a model 2 days ago
stabilityai/stable-point-aware-3d
liked a dataset 2 days ago
eltorio/ROCOv2-radiology
View all activity

Articles

Organizations

Hugging Face's profile picture Cohere For AI's profile picture SomosNLP's profile picture Libre Euro Lingua-Alliance's profile picture Hugging Face H4's profile picture Hugging Face OSS Metrics's profile picture Argilla's profile picture Blog-explorers's profile picture Hugging Face TB Research's profile picture ZeroGPU Explorers's profile picture h4-argilla-collab's profile picture mLLM multilingual's profile picture DIBT Spanish's profile picture Data is Better Together - Russian Language Team's profile picture Open Arabic LLM Leaderboard's profile picture Argilla Explorers's profile picture distilabel-internal-testing's profile picture Data Is Better Together's profile picture ORPO Explorers's profile picture Social Post Explorers's profile picture HuggingFaceFW-Dev's profile picture UCSF-JHU Opioid Industry Documents Archive's profile picture LLHF's profile picture SLLHF's profile picture Hugging Quants's profile picture argilla-internal-testing's profile picture Argilla Warehouse's profile picture rg-preview's profile picture Dataset Tools's profile picture open/ acc's profile picture Data Is Better Together Contributor's profile picture

dvilasuero's activity

New activity in CohereForAI/Global-MMLU about 1 month ago

Include argilla tag

2
#2 opened about 1 month ago by
dvilasuero
New activity in open-acc/README about 2 months ago
New activity in microsoft/orca-agentinstruct-1M-v1 about 2 months ago

message column is a str

3
#3 opened about 2 months ago by
dvilasuero
New activity in davidberenstein1957/vectorsearch-hub-datasets about 2 months ago

Typo search input

#1 opened about 2 months ago by
dvilasuero
New activity in GAIR/o1-journey 2 months ago
New activity in glaiveai/reflection-v1 3 months ago

Duplicates

3
#3 opened 3 months ago by
dvilasuero
New activity in dvilasuero/image-prefs 3 months ago
New activity in argilla/synthetic-data-generator 4 months ago
New activity in DIBT-Russian/prompt-translation-for-Russian 8 months ago

Upgrade Argilla server

1
#2 opened 8 months ago by
dvilasuero

Update Argilla version

#1 opened 8 months ago by
dvilasuero
New activity in dvilasuero/human-rights_config_space 9 months ago

Upload 2 files

#2 opened 9 months ago by
burtenshaw
New activity in abhishek/autotrain-llama3-orpo-v2 9 months ago

Adds dataset metadata

1
#1 opened 9 months ago by
dvilasuero