Please use this identifier to cite or link to this item: http://hdl.handle.net/1893/37587
Full metadata record
DC FieldValueLanguage
dc.contributor.authorZain, Noor Ulen_UK
dc.contributor.authorNaseem, Mohsin Razaen_UK
dc.contributor.authorAdeel, Ahsanen_UK
dc.contributor.editorCharpentier, Lucasen_UK
dc.contributor.editorChoshen, Leshemen_UK
dc.contributor.editorCotterell, Ryanen_UK
dc.contributor.editorGul, Mustafa Omeren_UK
dc.contributor.editorHu, Michael Yen_UK
dc.contributor.editorLiu, Jingen_UK
dc.contributor.editorJumelet, Jaapen_UK
dc.contributor.editorLinzen, Talen_UK
dc.contributor.editorMueller, Aaronen_UK
dc.contributor.editorRoss, Candaceen_UK
dc.contributor.editorShah, Raj Sanjayen_UK
dc.contributor.editorWarstadt, Alexen_UK
dc.contributor.editorWilcox, Ethan Gotlieben_UK
dc.contributor.editorWilliams, Adinaen_UK
dc.date.accessioned2025-11-26T01:12:47Z-
dc.date.available2025-11-26T01:12:47Z-
dc.date.issued2025en_UK
dc.identifier.urihttp://hdl.handle.net/1893/37587-
dc.description.abstractWe show that a tiny Co4 machine (CITATION) with a single layer, two heads, and 8M parameters, operating at O(N) computational cost (where N is the number of input tokens), in just 2 epochs outpaces GPT-2 (124M, 12 layers, O(N2)) and GPT-BERT (30M, 12 layers, O(N 2), both trained for 10 epochs. Co4 achieves orders-of-magnitude greater training efficiency on 10M tokens, demonstrating sample-efficient pretraining. On the BabyLM challenge evaluation pipeline, Co4 performs comparably or better across complex benchmarks, showing strong zero-shot and fine-tuning performance on SuperGLUE tasks. Specifically, Co4 outperforms GPT-2 in 5 out of 7 zero-shot metrics and 6 out of 7 fine-tuning tasks, and GPT-BERT in 4 out of 7 metrics in both cases. These results strongly suggest a need to rethink prevailing deep learning paradigms and associated scaling laws.en_UK
dc.language.isoenen_UK
dc.publisherAssociation for Computational Linguisticsen_UK
dc.relationZain NU, Naseem MR & Adeel A (2025) Single layer tiny Co4 outpaces {GPT}-2 and {GPT}-{BERT}. In: Charpentier L, Choshen L, Cotterell R, Gul MO, Hu MY, Liu J, Jumelet J, Linzen T, Mueller A, Ross C, Shah RS, Warstadt A, Wilcox EG & Williams A (eds.) <i>Proceedings of the First BabyLM Workshop</i>, volume Proceedings of the First BabyLM Workshop. Empirical Methods in Natural Language Processing, Hybrid, 04.11.2025. Association for Computational Linguistics, pp. 313-322. https://doi.org/10.18653/v1/2025.babylm-main.24en_UK
dc.rightsACL materials are Copyright © 1963–2025 ACL; Materials published in or after 2016 are licensed on a Creative Commons Attribution 4.0 International License.en_UK
dc.rights.urihttp://creativecommons.org/licenses/by/4.0/en_UK
dc.titleSingle layer tiny Co4 outpaces {GPT}-2 and {GPT}-{BERT}en_UK
dc.typeConference Paperen_UK
dc.identifier.doi10.18653/v1/2025.babylm-main.24en_UK
dc.identifier.pmid36568019en_UK
dc.citation.volumeProceedings of the First BabyLM Workshopen_UK
dc.citation.spage313en_UK
dc.citation.epage322en_UK
dc.citation.publicationstatusPublisheden_UK
dc.citation.peerreviewedRefereeden_UK
dc.type.statusVoR - Version of Recorden_UK
dc.contributor.funderAdvanced Research and Invention Agencyen_UK
dc.author.emailnoor.noorulzain@stir.ac.uken_UK
dc.citation.conferencedates2025-11-04en_UK
dc.citation.conferencelocationHybriden_UK
dc.citation.conferencenameEmpirical Methods in Natural Language Processingen_UK
dc.citation.date05/11/2025en_UK
dc.citation.isbnTODOen_UK
dc.contributor.affiliationComputing Science and Mathematics - Divisionen_UK
dc.contributor.affiliationComputing Scienceen_UK
dc.contributor.affiliationComputing Science and Mathematics - Divisionen_UK
dc.identifier.wtid2202837en_UK
dc.date.accepted2025-09-15en_UK
dcterms.dateAccepted2025-09-15en_UK
dc.date.filedepositdate2025-11-21en_UK
rioxxterms.apcpaiden_UK
rioxxterms.typeConference Paper/Proceeding/Abstracten_UK
rioxxterms.versionVoRen_UK
local.rioxx.authorZain, Noor Ul|en_UK
local.rioxx.authorNaseem, Mohsin Raza|en_UK
local.rioxx.authorAdeel, Ahsan|en_UK
local.rioxx.projectProject ID unknown|Advanced Research and Invention Agency|en_UK
local.rioxx.contributorCharpentier, Lucas|en_UK
local.rioxx.contributorChoshen, Leshem|en_UK
local.rioxx.contributorCotterell, Ryan|en_UK
local.rioxx.contributorGul, Mustafa Omer|en_UK
local.rioxx.contributorHu, Michael Y|en_UK
local.rioxx.contributorLiu, Jing|en_UK
local.rioxx.contributorJumelet, Jaap|en_UK
local.rioxx.contributorLinzen, Tal|en_UK
local.rioxx.contributorMueller, Aaron|en_UK
local.rioxx.contributorRoss, Candace|en_UK
local.rioxx.contributorShah, Raj Sanjay|en_UK
local.rioxx.contributorWarstadt, Alex|en_UK
local.rioxx.contributorWilcox, Ethan Gotlieb|en_UK
local.rioxx.contributorWilliams, Adina|en_UK
local.rioxx.freetoreaddate2025-11-21en_UK
local.rioxx.licencehttp://creativecommons.org/licenses/by/4.0/|2025-11-21|en_UK
local.rioxx.filename2025.babylm-main.24.pdfen_UK
local.rioxx.filecount1en_UK
local.rioxx.sourceTODOen_UK
Appears in Collections:Computing Science and Mathematics Conference Papers and Proceedings

Files in This Item:
File Description SizeFormat 
2025.babylm-main.24.pdfFulltext - Published Version657.69 kBAdobe PDFView/Open


This item is protected by original copyright



A file in this item is licensed under a Creative Commons License Creative Commons

Items in the Repository are protected by copyright, with all rights reserved, unless otherwise indicated.

The metadata of the records in the Repository are available under the CC0 public domain dedication: No Rights Reserved https://creativecommons.org/publicdomain/zero/1.0/

If you believe that any material held in STORRE infringes copyright, please contact library@stir.ac.uk providing details and we will remove the Work from public display in STORRE and investigate your claim.