https://aclanthology.org/2023.findings-acl.426/ ACL Logo ACL Anthology * FAQ(current) * Corrections(current) * Submissions(current) [ ] "Low-Resource" Text Classification: A Parameter-Free Classification Method with Compressors Zhiying Jiang, Matthew Yang, Mikhail Tsirlin, Raphael Tang, Yiqin Dai , Jimmy Lin --------------------------------------------------------------------- Abstract Deep neural networks (DNNs) are often used for text classification due to their high accuracy. However, DNNs can be computationally intensive, requiring millions of parameters and large amounts of labeled data, which can make them expensive to use, to optimize, and to transfer to out-of-distribution (OOD) cases in practice. In this paper, we propose a non-parametric alternative to DNNs that's easy, lightweight, and universal in text classification: a combination of a simple compressor like gzip with a k-nearest-neighbor classifier. Without any training parameters, our method achieves results that are competitive with non-pretrained deep learning methods on six in-distribution datasets.It even outperforms BERT on all five OOD datasets, including four low-resource languages. Our method also excels in the few-shot setting, where labeled data are too scarce to train DNNs effectively. Anthology ID: 2023.findings-acl.426 Volume: Findings of the Association for Computational Linguistics: ACL 2023 Month: July Year: 2023 Address: Toronto, Canada Venue: Findings SIG: Publisher: Association for Computational Linguistics Note: Pages: 6810-6828 Language: URL: https://aclanthology.org/2023.findings-acl.426 DOI: Bibkey: jiang-etal-2023-low Cite (ACL): Zhiying Jiang, Matthew Yang, Mikhail Tsirlin, Raphael Tang, Yiqin Dai, and Jimmy Lin. 2023. "Low-Resource" Text Classification: A Parameter-Free Classification Method with Compressors. In Findings of the Association for Computational Linguistics: ACL 2023, pages 6810-6828, Toronto, Canada. Association for Computational Linguistics. Cite (Informal): "Low-Resource" Text Classification: A Parameter-Free Classification Method with Compressors (Jiang et al., Findings 2023) Copy Citation: BibTeX Markdown MODS XML Endnote More options... PDF: https://aclanthology.org/2023.findings-acl.426.pdf PDF Cite Search --------------------------------------------------------------------- Export citation x * BibTeX * MODS XML * Endnote * Preformatted @inproceedings{jiang-etal-2023-low, title = "{``}Low-Resource{''} Text Classification: A Parameter-Free Classification Method with Compressors", author = "Jiang, Zhiying and Yang, Matthew and Tsirlin, Mikhail and Tang, Raphael and Dai, Yiqin and Lin, Jimmy", booktitle = "Findings of the Association for Computational Linguistics: ACL 2023", month = jul, year = "2023", address = "Toronto, Canada", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2023.findings-acl.426", pages = "6810--6828", abstract = "Deep neural networks (DNNs) are often used for text classification due to their high accuracy. However, DNNs can be computationally intensive, requiring millions of parameters and large amounts of labeled data, which can make them expensive to use, to optimize, and to transfer to out-of-distribution (OOD) cases in practice. In this paper, we propose a non-parametric alternative to DNNs that{'}s easy, lightweight, and universal in text classification: a combination of a simple compressor like \textit{gzip} with a $k$-nearest-neighbor classifier. Without any training parameters, our method achieves results that are competitive with non-pretrained deep learning methods on six in-distribution datasets.It even outperforms BERT on all five OOD datasets, including four low-resource languages. Our method also excels in the few-shot setting, where labeled data are too scarce to train DNNs effectively.", } Download as File Copy to Clipboard "Low-Resource" Text Classification: A Parameter-Free Classification Method with Compressors Zhiying Jiang author Matthew Yang author Mikhail Tsirlin author Raphael Tang author Yiqin Dai author Jimmy Lin author 2023-07 text Findings of the Association for Computational Linguistics: ACL 2023 Association for Computational Linguistics Toronto, Canada conference publication Deep neural networks (DNNs) are often used for text classification due to their high accuracy. However, DNNs can be computationally intensive, requiring millions of parameters and large amounts of labeled data, which can make them expensive to use, to optimize, and to transfer to out-of-distribution (OOD) cases in practice. In this paper, we propose a non-parametric alternative to DNNs that's easy, lightweight, and universal in text classification: a combination of a simple compressor like gzip with a k-nearest-neighbor classifier. Without any training parameters, our method achieves results that are competitive with non-pretrained deep learning methods on six in-distribution datasets.It even outperforms BERT on all five OOD datasets, including four low-resource languages. Our method also excels in the few-shot setting, where labeled data are too scarce to train DNNs effectively. jiang-etal-2023-low https://aclanthology.org/2023.findings-acl.426 2023-07 6810 6828 Download as File Copy to Clipboard %0 Conference Proceedings %T "Low-Resource" Text Classification: A Parameter-Free Classification Method with Compressors %A Jiang, Zhiying %A Yang, Matthew %A Tsirlin, Mikhail %A Tang, Raphael %A Dai, Yiqin %A Lin, Jimmy %S Findings of the Association for Computational Linguistics: ACL 2023 %D 2023 %8 July %I Association for Computational Linguistics %C Toronto, Canada %F jiang-etal-2023-low %X Deep neural networks (DNNs) are often used for text classification due to their high accuracy. However, DNNs can be computationally intensive, requiring millions of parameters and large amounts of labeled data, which can make them expensive to use, to optimize, and to transfer to out-of-distribution (OOD) cases in practice. In this paper, we propose a non-parametric alternative to DNNs that's easy, lightweight, and universal in text classification: a combination of a simple compressor like gzip with a k-nearest-neighbor classifier. Without any training parameters, our method achieves results that are competitive with non-pretrained deep learning methods on six in-distribution datasets.It even outperforms BERT on all five OOD datasets, including four low-resource languages. Our method also excels in the few-shot setting, where labeled data are too scarce to train DNNs effectively. %U https://aclanthology.org/2023.findings-acl.426 %P 6810-6828 Download as File Copy to Clipboard Markdown (Informal) ["Low-Resource" Text Classification: A Parameter-Free Classification Method with Compressors](https://aclanthology.org/ 2023.findings-acl.426) (Jiang et al., Findings 2023) * "Low-Resource" Text Classification: A Parameter-Free Classification Method with Compressors (Jiang et al., Findings 2023) ACL * Zhiying Jiang, Matthew Yang, Mikhail Tsirlin, Raphael Tang, Yiqin Dai, and Jimmy Lin. 2023. "Low-Resource" Text Classification: A Parameter-Free Classification Method with Compressors. In Findings of the Association for Computational Linguistics: ACL 2023, pages 6810-6828, Toronto, Canada. Association for Computational Linguistics. Copy Markdown to Clipboard Copy ACL to Clipboard Creative Commons License ACL materials are Copyright (c) 1963-2023 ACL; other materials are copyrighted by their respective copyright holders. Materials prior to 2016 here are licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 3.0 International License. Permission is granted to make copies for the purposes of teaching and research. Materials published in or after 2016 are licensed on a Creative Commons Attribution 4.0 International License. The ACL Anthology is managed and built by the ACL Anthology team of volunteers. Site last built on 13 July 2023 at 19:29 UTC with commit 59a0926.