Natalia Elvira

nataliaElv

AI & ML interests

Data curation, high-quality data, multilinguality, NLP & computational linguistics

Recent Activity

reacted to davanstrien's post with 🚀 1 day ago
The https://huggingface.co/datasets/data-is-better-together/fineweb-c dataset is growing! This week a few more languages have got 1,000 annotations for the educational quality of data from https://huggingface.co/datasets/HuggingFaceFW/fineweb-2. Why should you care? The quality of pre-training data can have a big impact on the performance of downstream language models trained on that data (https://huggingface.co/spaces/HuggingFaceFW/blogpost-fineweb-v1). Being able to filter by educational quality is on way of improving the quality of the data you use for training an LLM. Very importantly this approach can also reduce the amount of data needed for pertaining. Why not use an LLM? LLMs can be used to annotate educational quality for a subset of data. This data can then be used to train a smaller encoder only model to label the full dataset. However, this may not work well for languages outside of english. This is where fineweb-c (community) comes in. The community is annotating the educational quality of fineweb2 data. Currently 114 languages have some annotations. These annotations will enable a number of things: - Evaluate whether an LLM can label the educational quality for texts in that language well - Directly be used for training quality classifiers - Help discover other rules and huerisitcs for refining fineweb2 further for different languages. This week the following languages where done: Swedish thanks to: @Lauler @AntonVic @ohallstrom @bjarlestam @menbom @Ekgren @apsod Ukrainian thanks to: @hannayukhymenko @robinhad @realPivo @RabotiahovDmytro @reciprocate Assamese thanks to: @moyoor97 @Arpanjyoti @nawaf-helmi123 @pahigogoi1 @aelhence @kishorekashyap Want to learn more: https://huggingface.co/blog/davanstrien/fineweb2-community Contribute yourself here: https://huggingface.co/spaces/data-is-better-together/fineweb-c
View all activity

Articles

Organizations

Hugging Face's profile picture SomosNLP's profile picture Argilla's profile picture Blog-explorers's profile picture Argilla Explorers's profile picture Data Is Better Together's profile picture HuggingFaceFW-Dev's profile picture Hugging Face Discord Community's profile picture argilla-internal-testing's profile picture Argilla Warehouse's profile picture Dataset Tools's profile picture Coordination Nationale pour l'IA's profile picture Data Is Better Together Contributor's profile picture Bluesky Community's profile picture

nataliaElv's activity

New activity in bluesky-community/one-million-bluesky-posts about 2 months ago

Language tags

1
#1 opened about 2 months ago by
nataliaElv
New activity in nataliaElv/argilla-progress about 2 months ago

Update app.py

#1 opened about 2 months ago by
davidberenstein1957
New activity in huggingface-course/documentation-images about 2 months ago

More Argilla screenshots

#4 opened about 2 months ago by
nataliaElv

argilla-chapter-images

#3 opened about 2 months ago by
nataliaElv

Chapter 10 images

#2 opened about 2 months ago by
nataliaElv
New activity in nataliaElv/argilla about 2 months ago

Test

#1 opened about 2 months ago by
nataliaElv
New activity in somosnlp/somos-alpaca-es almost 2 years ago

Create guia-de-anotacion.md

2
#3 opened almost 2 years ago by
nataliaElv

Draft: create guia-de-anotacion.md

#2 opened almost 2 years ago by
nataliaElv