Óbudai Egyetem Digitális Archívum
    • magyar
    • English
  • magyar 
    • magyar
    • English
  • Bejelentkezés
Megtekintés 
  •   ÓDA repozitórium kezdőoldal
  • 5. Folyóiratcikkek
  • Acta Polytechnica Hungarica
  • 3. 2023
  • 3.6. 2023 Volume 20, Issue No. 5.
  • Megtekintés
  •   ÓDA repozitórium kezdőoldal
  • 5. Folyóiratcikkek
  • Acta Polytechnica Hungarica
  • 3. 2023
  • 3.6. 2023 Volume 20, Issue No. 5.
  • Megtekintés
JavaScript is disabled for your browser. Some features of this site may not work without it.

Training Experimental Language Models with Low Resources, for the Hungarian Language

Thumbnail
Megtekintés/Megnyitás
Yang_Varadi_134.pdf (467.4KB)
Metaadat
Teljes megjelenítés
Link a dokumentumra való hivatkozáshoz:
http://hdl.handle.net/20.500.14044/39245
Gyűjtemény
  • 3.6. 2023 Volume 20, Issue No. 5. [11]
Absztrakt
In recent years, natural language processing tasks, like sentiment analysis, can be solved with high performance techniques, if a pre-trained language model is fine-tuned. However, in most cases, the pre-training of language models require huge computational resources and training corpora. Our paper addresses the issue of developing deep neural network language models for low resourced languages, such as Hungarian. Pre-training language models like BERT, requires a prohibitive amount of computational power and huge amount of training data. Unfortunately, neither of these prerequisites are commonly available for low resource languages. The question is how well the system can perform with limited resources (both in data and hardware). We focus our research on five transformer models: ELECTRA, ELECTRIC, RoBERTa, BART and GPT-2. To evaluate our models, we fine-tuned the models in six different natural language processing tasks: sentence-level sentiment analysis, named entity recognition, noun phrase chunking, extractive summarization and abstractive summarization. Our results suggest that while our experimental models obviously cannot surpass the performance of the state-of-the-art Hungarian BERT model, they require a smaller carbon footprint, may bring neural network technology to mobile applications and, finally, they may lower the threshold to engaging with neural network technology in low resourced languages, which has been an obstacle so far, in the synergistic co-development of cognitive info-communication systems and its related disciplines.
Cím és alcím
Training Experimental Language Models with Low Resources, for the Hungarian Language
Szerző
Yang, Zijian Győző
Váradi, Tamás
Megjelenés ideje
2023
Hozzáférés szintje
Open access
ISSN, e-ISSN
1785-8860
Nyelv
en
Terjedelem
20 p.
Tárgyszó
ELECTRA, ELECTRIC, RoBERTa, BART, GPT-2, sentiment analysis, named entity recognition, noun phrase chunking, text summarization
Változat
Kiadói változat
Egyéb azonosítók
DOI: 10.12700/APH.20.5.2023.5.11
A cikket/könyvrészletet tartalmazó dokumentum címe
Acta Polytechnica Hungarica
A forrás folyóirat éve
2023
A forrás folyóirat évfolyama
20. évf.
A forrás folyóirat száma
5. sz.
Műfaj
Tudományos cikk
Tudományterület
Műszaki tudományok - informatikai tudományok
Egyetem
Óbudai Egyetem

DSpace software copyright © 2002-2016  DuraSpace
Kapcsolat | Visszajelzés
Theme by 
Atmire NV
 

 

Böngészés

A teljes ÓDA-banKategóriák és gyűjteményekMegjelenés dátumaSzerzőCímTárgyszóA gyűjteménybenMegjelenés dátumaSzerzőCímTárgyszó

Személyes felhasználói fiók

BejelentkezésRegisztráció

DSpace software copyright © 2002-2016  DuraSpace
Kapcsolat | Visszajelzés
Theme by 
Atmire NV