24 5 6

Francisco Aranda

frascuchon

AI & ML interests

None yet

Recent Activity

reacted to davanstrien's post with 🚀 1 day ago

The https://huggingface.co/datasets/data-is-better-together/fineweb-c dataset is growing! This week a few more languages have got 1,000 annotations for the educational quality of data from https://huggingface.co/datasets/HuggingFaceFW/fineweb-2. Why should you care? The quality of pre-training data can have a big impact on the performance of downstream language models trained on that data (https://huggingface.co/spaces/HuggingFaceFW/blogpost-fineweb-v1). Being able to filter by educational quality is on way of improving the quality of the data you use for training an LLM. Very importantly this approach can also reduce the amount of data needed for pertaining. Why not use an LLM? LLMs can be used to annotate educational quality for a subset of data. This data can then be used to train a smaller encoder only model to label the full dataset. However, this may not work well for languages outside of english. This is where fineweb-c (community) comes in. The community is annotating the educational quality of fineweb2 data. Currently 114 languages have some annotations. These annotations will enable a number of things: - Evaluate whether an LLM can label the educational quality for texts in that language well - Directly be used for training quality classifiers - Help discover other rules and huerisitcs for refining fineweb2 further for different languages. This week the following languages where done: Swedish thanks to: @Lauler @AntonVic @ohallstrom @bjarlestam @menbom @Ekgren @apsod Ukrainian thanks to: @hannayukhymenko @robinhad @realPivo @RabotiahovDmytro @reciprocate Assamese thanks to: @moyoor97 @Arpanjyoti @nawaf-helmi123 @pahigogoi1 @aelhence @kishorekashyap Want to learn more: https://huggingface.co/blog/davanstrien/fineweb2-community Contribute yourself here: https://huggingface.co/spaces/data-is-better-together/fineweb-c

upvoted a paper 1 day ago

Human Still Wins over LLM: An Empirical Study of Active Learning on Domain-Specific Annotation Tasks

liked a Space 2 days ago

argilla/synthetic-data-generator

View all activity

Articles

How to optimize your data labelling project with custom interfaces

Oct 16, 2024

• 18

Organizations

frascuchon's activity

reacted to davanstrien's post with 🚀 1 day ago

Post

1127

The data-is-better-together/fineweb-c dataset is growing!

This week a few more languages have got 1,000 annotations for the educational quality of data from HuggingFaceFW/fineweb-2.

Why should you care?

The quality of pre-training data can have a big impact on the performance of downstream language models trained on that data ( HuggingFaceFW/blogpost-fineweb-v1).

Being able to filter by educational quality is on way of improving the quality of the data you use for training an LLM. Very importantly this approach can also reduce the amount of data needed for pertaining.

Why not use an LLM?

LLMs can be used to annotate educational quality for a subset of data. This data can then be used to train a smaller encoder only model to label the full dataset. However, this may not work well for languages outside of english. This is where fineweb-c (community) comes in.

The community is annotating the educational quality of fineweb2 data. Currently 114 languages have some annotations. These annotations will enable a number of things:

- Evaluate whether an LLM can label the educational quality for texts in that language well
- Directly be used for training quality classifiers
- Help discover other rules and huerisitcs for refining fineweb2 further for different languages.

This week the following languages where done:

Swedish thanks to: @Lauler @AntonVic @ohallstrom @bjarlestam @menbom @Ekgren @apsod

Ukrainian thanks to: @hannayukhymenko @robinhad @realPivo @RabotiahovDmytro @reciprocate

Assamese thanks to: @moyoor97 @Arpanjyoti @nawaf-helmi123 @pahigogoi1 @aelhence @kishorekashyap

Want to learn more: https://huggingface.co/blog/davanstrien/fineweb2-community

Contribute yourself here: data-is-better-together/fineweb-c

1 reply

upvoted a paper 1 day ago

Human Still Wins over LLM: An Empirical Study of Active Learning on Domain-Specific Annotation Tasks

Paper • 2311.09825 • Published Nov 16, 2023 • 1

liked 2 Spaces 2 days ago

Running

352

🧬

Synthetic Data Generator

Build datasets using natural language

Running

✍️✨

Dataset ReWriter

ReWrite datasets with a text instruction

updated a dataset 24 days ago

frascuchon/new_finepersonas-v0.1-tiny-flux-schnell

Viewer • Updated 24 days ago • 100 • 101

updated a dataset 26 days ago

frascuchon/awesome-chatgpt-prompts

Viewer • Updated 26 days ago • 170 • 117

New activity in frascuchon/awesome-chatgpt-prompts 26 days ago

Add responses for user frascuchon

#3 opened 26 days ago by

frascuchon

Add responses for user frascuchon

#4 opened 26 days ago by

frascuchon

Add responses for user frascuchon

#2 opened 26 days ago by

frascuchon

Add responses for user frascuchon

#1 opened 26 days ago by

frascuchon

updated 2 datasets 26 days ago

frascuchon/jfcalvo_finepersonas-with-all-questions

Viewer • Updated 26 days ago • 350 • 25

frascuchon/hydrate-test

Viewer • Updated 26 days ago • 10k • 36

updated 4 Spaces about 1 month ago

Running on CPU Upgrade

🌐

FineWeb-c - Annotation

New activity in data-is-better-together/fineweb-c about 1 month ago

Configure space with the progress-sharing feature

#1 opened about 1 month ago by

frascuchon

posted an update about 1 month ago

Post

386

🚀 Argilla v2.5.0 is out! 🎉
We’re excited to announce the latest version of Argilla, packed with features to make your data annotation workflows more powerful and seamless. Here’s what’s new:

✨ 1. Argilla Webhooks
With Argilla webhooks, you can:
* Trigger custom workflows
* Seamlessly integrate with external tools
* Build custom event-driven pipelines

🐍 2. Support for Python 3.13 and Pydantic v2
Argilla v2.5.0 now runs on:
* Python 3.13 for enhanced compatibility and speed
* Pydantic v2 for improved performance and type validation

🎨 3. Redesigned Home Page
Argilla's home page has been redesigned to provide a better user experience, showing a new dataset card view, which provides a better overview of the datasets and annotation progress.

📖 Read the full release notes 👉 https://github.com/argilla-io/argilla/releases/tag/v2.5.0)
⬇️ Update now 👉 https://pypi.org/project/argilla)
or use the live demo 👉 argilla/argilla-template-space