https://knowingmachines.org/critical-field-guide-mob A B C D E F G H I J K L M N O P Q R S T U V W 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 A CRITICAL FIELD GUIDE FOR WORKING WITH MACHINE LEARNING DATASETS Written by Sarah Ciston {1} Editors: Mike Ananny {2} and Kate Crawford {3} 20 Part of the Knowing Machines research project. 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 TABLE OF CONTENTS 40 1. Introduction to Machine Learning Datasets 41 2. Benefits: Why Approach Datasets Critically? 42 3. Parts of a Dataset 43 4. Types of Datasets 44 5. Transforming Datasets 45 6. The Dataset Lifecycle 46 7. Cautions & Reflections from the Field 47 8. Conclusion 48 49 50 1 51 INTRODUCTION TO MACHINE LEARNING DATASETS 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 Maybe you're an engineer creating a new machine vision system to track birds. You might be a journalist using social media data to research Costa Rican households. You could be a researcher who stumbled upon your university's archive of handwritten 72 census cards from 1939. Or a designer creating a chatbot that relies on large language models like GPT-3. Perhaps you're an artist experimenting with visual style combinations using DALLE-2. Or maybe you're an activist with an urgent story that needs telling, and you're searching for the right dataset to tell it. 73 WELCOME. No matter what kind of datasets you're using or want to use, whether you're curious but intimidated by machine learning or 74 already comfortable, this work is complicated. Because machine learning relies on datasets, and because datasets are always tangled up in the ways they're created and used, things can get messy. You may have questions like: 75 76 Does this dataset tell the story of my research in the way I want? 77 How do the dataset pre-processing methods I choose affect my outcomes? 78 How might this dataset contribute to creating errors or causing harm? 79 More than likely you will encounter at least some of these conundrums -- as many of us who work with machine learning 80 datasets do. Anyone using datasets will weigh choices and make tradeoffs. There are no universal answers and no perfect actions -- just a tangle of dataset forms, formats, relationships, behaviors, histories, intentions, and contexts. When choosing and using machine learning datasets, how do you 81 deal with the issues they bring? How can you navigate the mess thoughtfully and intentionally? Let's jump in. 82 83 84 INTRODUCTION TO MACHINE LEARNING DATASETS 85 1.1 86 WHAT IS THIS GUIDE ? 87 88 89 Machine learning datasets are powerful but unwieldy. They are often far too large to check all the data manually, to look for inaccurate labels, dehumanizing images, or other widespread issues. Despite the fact that datasets commonly contain problematic material -- whether from a technical, legal, or ethical perspective -- datasets are also valuable resources when handled carefully and critically. This guide offers questions, 90 suggestions, strategies, and resources to help people work with existing machine learning datasets at every phase of their lifecycle. Equipped with this understanding, researchers and developers will be more capable of avoiding the problems unique to datasets. They will also be able to construct more reliable, robust solutions, or even explore promising new ways of thinking with machine learning datasets that are more critical and conscientious. {4} , {5} 91 If you aren't sure whether this guide is for you, consider the 92 many places you might find yourself working with machine learning datasets. This guide can be helpful if you are... 93 94 - making a model 95 - working with a pre-trained model 96 - researching an existing machine learning tool 97 - teaching with datasets 98 - creating an index or inventory 99 - concerned about how datasets describe you or your community 100 - learning about datasets by exploring one 101 - stewarding or archiving datasets 102 - investigating as an artist, activist, or developer 103 104 105 This list is non-exhaustive, of course. Datasets are 106 being used widely across countless domains and 107 industries. How else can you imagine working with 108 machine learning datasets? 109 110 111 The appetite for massive datasets is huge and still accelerating, fueled by the perceived promise of machine learning to convert data into meaningful, monetizable information.{6} Too often, this work is done without regard for how datasets can be partial, imperfect, and historically skewed. Take the widely publicized examples of police departments and courts selecting "future criminals" from software that relied on historical crime records, which ProPublica journalists found was grossly inaccurate and targeted Black people in its predictions 112 [6]. More troubling still, researchers (and public and private organizations) continue to make use of such datasets despite learning of their harms -- perhaps because they seem more efficient or effective, because they are already part of common practices in their communities, or simply because they are the most readily available options. This is exactly why critical care is so needed -- datasets' potential harms are subtle, localized, and complex. You will need to make conscientious decisions and compromises when working with any dataset. There is no perfect representation, no correct procedure, and no ideal dataset. This guide aims to help you navigate the complexity of working with datasets, giving you ways to approach conundrums carefully and thoughtfully. Section 1 describes how DATA and DATASETS are dynamic research materials, and Section 2 outlines the BENEFITS of working critically with datasets. Then you'll find more on 113 the common PARTS of datasets (Section 3), examples of the TYPES of datasets you may encounter (Section 5), and how to TRANSFORM datasets (Section 4) -- all to help make critical choices easier. Then Section 6 provides a DATASET LIFECYCLE framework with important questions to ask as you engage critically at each stage of your work. Finally, Section 7 offers some CAUTIONS & REFLECTIONS for careful dataset stewardship. 114 115 FIELD GUIDES AND DATASETS AS FORMS 116 117 118 119 120 121 122 123 124 The field guide format frames this text because, like 125 datasets, field guides teach their readers particular 126 ways of looking at the world -- for better and for worse. 127 Carried in a knapsack, a birder might use their field 128 guide to confirm a species sighting in a visual index. A 129 hiker might read trail warnings to prepare for their 130 trek. With these practical uses, the field guide speaks 131 to a desire to connect deeply with dataset tools and 132 practices, and a sense of careful responsibility that 133 data stewardship shares with environmental stewardship. 134 However, naturalist field guides also draw on the same 135 problematic histories of classifying and organizing 136 information that are foundational to many machine 137 learning tasks today. This critical field guide aims to 138 help bring understanding to the complexities of 139 datasets, so that the decisions you make while using 140 them are conscientious. It invites you to mess with 141 these messy forms and to approach any logic of 142 classification with a critical eye. 143 144 145 146 147 148 149 150 151 152 153 154 155 156 When we say CLASSIFICATION in this guide, generally we 157 refer to the choices, logics, and paradigms that inform 158 sociotechnical communities of practice -- how people sort 159 and are sorted into categories, how those categories 160 come to be seen as dominant and naturalized, and how 161 people are differently affected by those categories. We 162 acknowledge that the term CLASSIFICATION also refers to 163 specific machine learning tasks that label and sort 164 items in a dataset by discrete categories. For example, 165 asking whether an image is a dog or a cat is handled by 166 a classification task. These are distinguished from 167 REGRESSION tasks, which show the relationship between 168 features in a dataset, for example sorting dogs by their 169 age and number of spots. In this guide, we will specify 170 'tasks' when referring to these techniques, but simply 171 say 'classification' when referring to the 172 sociotechnical phenomenon more broadly. {7} 173 174 175 176 177 178 179 180 181 INTRODUCTION TO MACHINE LEARNING DATASETS 182 1.2 183 WHAT ARE DATA? 184 185 186 187 188 DATA ARE CONTINGENT ON HOW WE USE THEM 189 DATA are values assigned to any 'thing', and the term can be applied to almost anything. Numbers, of course, can be data; but so can emails, a collection of scanned manuscripts, the steps you walked to the train, the pose of a dancer, or the breach of a whale. How you think about the information is what makes it data. Philosopher of science Sabina Leonelli sees data as a 190 "relational category" meaning that, "What counts as data depends on who uses them, how, and for which purposes." She argues data are "any product of research activities [...] that is collected, stored, and disseminated in order to be used as evidence for knowledge claims" [8]. This definition reframes data as contingent on the people who create, use, and interact with them in context. 191 192 193 DATA MUST BE MADE, AND MAKING DATA SHAPES DATA 194 As a form of information, {8} data do not just exist but have to be generated, through collection by sensors or human effort. Sensing, observing, and collecting are all acts of interpretation that have contexts, which shape the data. For example, when collecting images of faces using infrared cameras, that data can provide heat signatures but not the eye color of its subjects. Studies are designed with specific equipment to achieve their goals and not others. Whether they are 195 quantitative data captured with a sensor or qualitative data described in an interview, the context in which that data is collected has already created a limit for what it can represent and how it can be used. It is easy to think that calling information "data" makes it discrete, separate, fixed, organized, computable -- static [1]. But dataset users impose these qualities on information temporarily -- to organize it into familiar forms that suit machine learning tasks and other algorithmic systems. 196 197 198 MACHINE LEARNING, DEEP LEARNING, NEURAL NET, ALGORITHM, MODEL -- WHAT'S THE DIFFERENCE? 199 200 201 202 203 204 205 206 An ALGORITHM is a set of instructions for a procedure, 207 whether in the context of machine learning or another 208 task. Algorithms are often written in code for machines 209 to process, but they are also widely used in any system 210 of step-by-step instructions (e.g. cooking recipes). 211 Algorithms are not a modern Western invention, but 212 predate computation by thousands of years, as technology 213 culture researcher Ted Striphas has shown [16]. That 214 said, algorithms stayed associated mainly with 215 mathematical calculation until quite recently, according 216 to historian of science Lorraine Datson, who traces 217 their expansion into a computational catch-all in the 218 mid-20th century [17]. 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 A MODEL is the result of a machine learning algorithm, 236 once it includes revisions that take into account the 237 data it was exposed to during its training. It is the 238 saved output of the training process, ready to make 239 predictions about new data. One way to think of a model 240 is as a very complex mathematical formula containing 241 millions or billions of variables (values that can 242 change). These variables, also called model parameters, 243 are designed to transform a numerical input into the 244 desired outputs. The process of model training entails 245 adjusting the variables that make up the formula until 246 its output matches the desired output. 247 248 Much focus is put on machine learning models, but models 249 depend directly on datasets for their predictions. While 250 a model is not part of a dataset, it is deeply shaped by 251 the datasets it is based upon. Traces of those datasets 252 remain embedded within the model no matter how it is 253 used next. (This guide won't cover the detailed aspects 254 of working critically with machine learning models and 255 understanding how they learn -- that's a whole other 256 discussion. Terms like 'activation functions', 'loss 257 functions', 'learning rates', and 'fine-tuning' give a 258 taste of the many human-guided processes behind model 259 making, an active conversation and set of practices that 260 are beyond the scope of this guide.) 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 Artificial NEURAL NETWORKS describe some of the ways to 277 structure machine learning models (see TYPES inSection 4 278 ), including making large language models. Named for the 279 inspiration they take from brain neurons (very 280 simplified), they move information through a series of 281 nodes (steps) organized in layers or sets. Each node 282 receives the output of the previous layers' nodes, 283 combines them using a mathematical formula, then passes 284 the output to the next layer of nodes. 285 286 287 288 289 290 291 292 293 294 MACHINE LEARNING is a set of tools used by computer programmers to find a formula that best describes (or 295 models) a dataset. Whereas in other kinds of software the programmer will write explicit instructions for 296 every part of a task, in machine learning, programmers will instruct the software to adjust its code based on 297 the data it processes, thus "learning" from new information [ 18] . Its learning is unlike human 298 understanding and the term is used metaphorically. 299 Some formulas are "deeper" than others, so called because they contain many more variables, and DEEP 300 LEARNING refers to the use of complex, many layers in a machine learning model. Due to their increasing 301 complexity, the outputs of machine learning models are not reliable for making decisions about people, 302 especially in highly consequential cases. When working with datasets, include machine learning as one suite of 303 options in a broader toolkit -- rather than a generalizable multi-tool for every task. 304 305 306 307 308 309 310 INTRODUCTION TO MACHINE LEARNING DATASETS 311 1.3 312 WHAT ARE DATASETS? 313 314 315 A DATASET can be any kind of collected, curated, interrelated data. Often, datasets refer to large collections of data used in computation, and especially in machine learning. Information 316 collections are transformed into datasets through a LIFECYCLE of processes (collection/selection, cleaning and analyzing, sharing and deprecating), which shape how that information is understood. (For critical questions you can ask at each phase of a dataset's lifecycle, see Section 6.) 317 318 319 DATASETS ARE TIED TO THEIR MAKERS 320 The many choices that go into dataset creation and use make them extremely dynamic. They always reflect the circumstances of their making -- the constraints of tools, who wields them and how, even who can afford the equipment to train, store, and transmit data. For example, analysis of national datasets in the Processing Citizenship project revealed how some European nations collected information differently, with a range of 321 specificity in categories like 'education level' or 'marital status'. Software engineer Wouter Van Rossem and science and technology studies professor Annalisa Pelizza examined not only the data in that dataset, but how they were labeled, organized, and utilized to show that these reflected how nations perceived the migrants they cataloged [ 19] . When gathered by a different group, using different tools, a dataset will be quite different -- even if it attempts to collect similar information. 322 323 DATASETS ARE TIED TO THEIR SETTINGS 324 Datasets can be frustratingly limited, but this does not mean they are static; instead, the information in datasets is always wrapped up in the contexts that make and use them. Media scholar Yanni Alexander Loukissas, author of All Data Are Local, calls datasets "data settings," arguing that "data are indexes to local knowledge." They remain tied to the communities, 325 individuals, organisms, and environments where they were created. Instead of treating data as independent authorities, he says we should ask, "Where do data direct us, and who might help us understand their origins as well as their sites of potential impact?" [1]. These questions extend the possibilities for exploring datasets as dynamic materials. Therefore, datasets must be used carefully, with consideration for their material connection to their origins. 326 327 328 329 For your consideration: How does framing information as 330 "data" change your relationship to it? What other forms 331 of information do you work with? What kinds of 332 information should not be included in datasets? 333 334 335 336 337 338 2 339 BENEFITS: WHY APPROACH DATASETS CRITICALLY? 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 HERE ARE SOME EXAMPLES OF HOW DATASET STEWARDSHIP CAN BENEFIT YOUR PRACTICE, AS WELL AS BENEFIT OTHERS: 361 362 363 MORE ROBUST DATASETS ARISE BY CONSIDERING MULTIPLE PERSPECTIVES AND WORKING TO REDUCE BIAS. 364 365 366 TO-DO: Include interdisciplinary, intersectional 367 communities in designing, developing, implementing, and evaluating your work. (See ALTERNATIVE APPROACHES TO DATASET PRACTICES) 368 369 370 371 MORE RELIABLE RESULTS COME FROM ANTICIPATING AND ADDRESSING CONTINGENCIES LIKE DEPRECATED DATASETS AND UNINFORMED CONSENT. 372 373 374 TO-DO: Apply checkpoints at each stage, asking critical 375 questions about data provenance and reflecting on your own methodologies. 376 377 378 GAIN INCREASED PROTECTION FROM LIABILITY FOR DATASETS WITH LEGAL 379 OR ETHICAL ISSUES BY PROACTIVELY ADDRESSING POTENTIAL CONCERNS BEFORE USE. 380 381 382 TO-DO: This does not constitute legal advice. However, always perform due diligence before working with existing datasets, including checking any licenses or terms of use. Simply downloading some datasets can create legal liability [20] , [21]. So try to be aware 383 of potential consent issues, misuse, or ethical concerns beyond those outlined by the dataset creators, especially as they may have changed since creation or arise from your new usage. You can check data repositories and data journalism to see how datasets have already been used. 384 385 386 387 CRITICAL PRACTICES ARE BECOMING FIELD-STANDARD AND REQUIRED FOR ACCESS TO TOP CONFERENCES AND JOURNALS. 388 389 390 TO-DO: Help shape the future of the field by modeling 391 and advocating for best practices. Suggest new frameworks and methods for making, using, and deprecating datasets. 392 393 394 395 MORE CAREFUL AND CONSCIENTIOUS OUTCOMES FOR THOSE IMPACTED BY RESULTS. 396 397 398 TO-DO: Engage the people and groups affected by datasets 399 and your use of them, to learn what careful and conscientious practices mean to them. 400 401 402 OPEN-SOURCE, OPEN-ACCESS, AND OPEN RESEARCH COMMUNITIES BUILD 403 POSITIVE FEEDBACK LOOPS THROUGH DATASET STEWARDSHIP OF RELIABLE MATERIALS. 404 405 406 407 TO-DO: Share datasets responsibly, through centralized repositories and with thorough documentation. 408 409 410 NO NEUTRAL CHOICE (OR NON-CHOICE) EXISTS. "WHEN THE FIELD OF AI 411 BELIEVES IT IS NEUTRAL," SAYS AI RESEARCHER PRATYUSHA KALLURI, IT "BUILDS SYSTEMS THAT SANCTIFY THE STATUS QUO AND ADVANCE THE INTERESTS OF THE POWERFUL" [23]. 412 413 414 TO-DO: Working with datasets brings challenges that need conversations and multiple perspectives. Discuss issues 415 with your team using the Dataset's Lifecycle questions in Section 6, plus the wide range of critical positions shared in the "Critical Dataset Studies Reading List" compiled by the Knowing Machines research project [22]. Before publishing or launching your work, ask hard questions and share your project with informal readers 416 within your networks who can provide constructive feedback. Go slow. Pause or even stop a project if needed. Remember that taking "no position" on a dataset's 417 ethical questions is still taking a position. Consider the tradeoffs for choosing one dataset or technique over another. 418 419 420 421 3 422 PARTS OF A DATASET 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 What actually makes up a machine learning dataset, practically 446 speaking? Here are some of the key terms that are helpful for understanding their parts and dynamics: 447 448 449 INSTANCE 450 One data point being processed or sorted, often viewed as a row in a table. For example, in a training dataset for a classification task that will sort images of dogs from cats, one 451 instance might include the image of a dog and the label "dog," while another instance would be an image of a cat and the label "cat" as well as other pertinent metadata (see also LABEL, METADATA, and TRAINING DATA below in this section and SUPERVISED machine learning in Section 4). 452 453 454 FEATURE 455 One attribute being analyzed, considered, or explored across the dataset, often viewed as a column in a table. Features can be any machine-readable (i.e. numeric) form of an instance: images converted into a sequence of pixels, for example. Note: 456 Researchers often select and "extract" the features most relevant for their purpose. Features are not given by default. They are the results of decisions made by datasets' creators and users. (For more discussion of ENGINEERING FEATURES see Section 5.) 457 458 459 LABEL 460 The results or output assigned by a machine learning model, or a descriptor included in a training dataset meant for the model to 461 practice on as it is built, or in a testing or benchmark dataset used for evaluation or verification. (See Section 7.2.2 for more on labels' creation and their potentially harmful impacts.) 462 463 464 METADATA 465 Data about data, metadata is supplementary information that describes a file or accompanies other content, e.g. an image from your camera comes with the date and location it was shot, lens aperture, and shutter speed. Metadata can describe 466 attributes of content and can also include who created it, with what tools, when, and how. Metadata may appear as TABULAR DATA (a table) and can include captions, file names and sizes, catalog index numbers, or almost anything else. Metadata are often subject- or domain-specific, reflecting how a group organizes, standardizes and represents information [23]. 467 468 469 DATASHEET 470 A document describing a dataset's characteristics and composition, motivation and collection processes, recommended usage and ethical considerations, and any other information to help people choose the best dataset for their task. Datasheets were proposed by diversity advocate and computer scientist 471 Timnit Gebru, et al., as a field-wide practice to "encourage reflection on the process of creating, distributing, and maintaining a dataset, including any underlying assumptions, potential risks or harms, and implications for use" [24]. Datasheets are also resources to help people select and adapt datasets for new contexts. 472 473 474 SAMPLE 475 A selection of the total dataset, whether chosen at random or using a particular feature or property; samples can be used to 476 analyze a dataset, perform testing, or train a model. For more on practices like sampling that transform datasets, see Section 5. 477 478 479 TRAINING DATA 480 A portion of the full dataset used to create a machine learning model, which will be kept out of later testing phases. Imagining a model like a student studying for exams, you could liken the training data to their study guide which they use to practice 481 the material. For example, in supervised machine learning (see Section 4), training data includes results like those the model will be asked to generate, e.g. labeled images. Training datasets can never be neutral, and they commonly "inherit learned logic from earlier examples and then give rise to subsequent ones," says critical AI scholar Kate Crawford [25] . 482 483 484 VALIDATION DATA 485 A portion of the full dataset that is separated from training data and testing data, validation data is held back and used to compare the performance of different design details. Validation data is separate from testing data, because validation data is 486 used during the training process to optimize the model while adjustments are being made; therefore, the resulting model will be familiar with its data. That means separate testing data is still needed to confirm how the final model performs. Imagine validation data as practice tests that programmers can administer to check on the model's progress so far. 487 488 489 TESTING DATA 490 A portion of the full dataset that is separated from the training data and validation data, and that is not involved in 491 creation of a machine learning model. Testing data is then run through the completed model in order to assess how well it functions. Testing data for the model would be similar to the student's final exam. 492 493 494 TENSORS: SCALARS, VECTORS, MATRICES (oh my!) 495 Software for working with machine learning datasets organizes information in numerical relationships, in grids called TENSORS. Understanding tensors can help you understand how data are viewed, compared, and manipulated in computational models. Their 496 grids can have many dimensions, not only two-dimensional X-and-Y graphs [26] . A SCALAR describes a single number. A VECTOR is a list (aka an array), like a line of numbers. A MATRIX is a 2D tensor, like a rectangle of numbers. And a grid of three (or more) dimensions is a TENSOR, like a cube of numbers, or a many-dimensional cube of numbers. 497 498 499 DATA SUBJECTS 500 The people and other beings whose data are gathered into a 501 dataset. Even if identifying information has been removed, datasets are still connected to the subjects they claim to represent. 502 503 504 DATA SUBJECTEES 505 This new and somewhat unwieldy term is used here to describe people impacted directly or indirectly by datasets, distinct from data subjects. Data subjectees include anyone affected by predictions made with machine learning models, for example 506 someone forced to use a facial detection system to board a flight or eye-tracking software to take a test at school. Similarly, Batya Friedman and David G. Hendry of the Value Sensitive Design Lab distinguish between "direct" and "indirect stakeholders" to describe the different types of entanglement with technologies [ 27] . 507 508 509 510 511 For your consideration: What other parts of a dataset 512 are not included here but could be? How do you see 513 dataset parts differently when you consider them within 514 "data settings," or contexts, tied to data subjects and 515 data subjectees? [1] What kinds of contexts are 516 impossible to include in datasets? 517 518 519 520 521 522 4 523 TYPES OF DATASETS 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 TYPES OF DATASETS 547 4.1 548 WHAT DISTINGUISHES TYPES OF DATASETS? 549 550 551 552 You may choose a dataset based on what it contains, how it is formatted, or other needs. For example, computer vision datasets include thousands of IMAGE or VIDEO files, while natural language processing datasets contain millions of bytes of TEXT. You may work with waveforms as SOUND files or time series data, 553 or network GRAPH stored in structured text formats like JSON. In tables you might find PLACE data as geographic coordinates or X-Y-Z coordinates, TIME as historical date sequences or milliseconds. Likely, you'll work with other types, too, or with combinations of MULTIMODAL data. Each dataset may include corresponding METADATA, documentation, and (hopefully) a complete DATASHEET. 554 555 You can also consider datasets based on whether the information is STRUCTURED, such as tabular data formatted in a table with labeled columns, or UNSTRUCTURED, such as plain text files or 556 unannotated images. Annotating or coding a dataset prepares it for analysis, including supervised machine learning; and annotation raises important questions about labor, classification, and power. (See Section 6.1 for more on annotation and labeling.) 557 558 Datasets for SUPERVISED machine learning need to include labels for at least a portion of the data that the system is designed to "learn." This means, for example, that a dataset for object 559 recognition would contain images as well as a table to describe the manually located object(s) they contain. It might have columns for the object name or label, as well as coordinates for the object position or outline, and the corresponding image's file name or index number. 560 561 In contrast, UNSUPERVISED machine learning looks for patterns that are not yet labeled in the dataset. It uses different kinds of machine learning algorithms, such as clustering groups of data together using features they share. However, it would be a misnomer to think that conclusions drawn from unsupervised machine learning are somehow more pure or rational. Much human judgment goes into developing an unsupervised machine learning model -- from adjusting weights and parameters to comparing 562 models' performance. Often supervised and unsupervised approaches are used in combination to ask different kinds of questions about the dataset. Other kinds of machine learning approaches (like reinforcement learning) don't fall neatly into these high-level categories. (For a discussion of deprecated datasets, see Section 7.2.3, and for critical questions at every stage of working with datasets, see Section 6.) 563 564 565 566 TYPES OF DATASETS 567 4.2 568 EXAMPLES: HOW HAVE RESEARCHERS, ENGINEERS, JOURNALISTS, AND ARTISTS PUT DATASETS TO USE? 569 570 571 When starting a project, you may not know what kind of dataset you need. You might work with a particular kind of media or file type most often, so you start there -- or maybe you want to try a 572 new form. You may start with a curiosity, and you're open to datasets in any format, from any source. To spark your imagination, here are four projects that used pre-existing datasets in novel and creative ways: 573 574 JOURNALISTS UNCOVER RAINFOREST EXPLOITATION WITH GEOSPATIAL DATA Brazilian investigative journalists at Armando.info, collaborating with El Pais and Earthrise Media, used field reports and satellite images to find deforestation, hidden runways, and illegal mining in the Venezuelan and Brazilian Amazon. Through computer vision analysis developed from analog maps and used on imagery from a European Space Agency satellite, the journalists compared this analysis with existing information, including complaints from Indigenous communities. 575 "It's not that this was a technology-only job," says Joseph Poliszuck, Armando.info's co-founder. "The technology allowed us to go into the field without being blindfolded" [ 28] . Using similar methods, Pulitzer fellow Hyury Potter detected approximately 1,300 illegal runways, more than the number of legally registered ones in the Brazilian Amazon. Combining data work with fieldwork in international collaborations helped these journalists connect local stories to larger scale climate crises and to support communities' efforts to create change. 576 577 578 HISTORIANS ASSEMBLE FRAGMENTS OF ANCIENT TEXTS Researchers from Google's DeepMind used a neural net on an existing scholarly dataset to complete, date, and find attributions for fragments of ancient texts. They drew on 178,551 ancient inscriptions written on stone, pottery, metal, and more. that had been transcribed in the text archive Packard Humanities Institute's Searchable Greek Inscriptions [29] . They 579 said that the "process required rendering the text machine-actionable [in plain text formats], normalizing epigraphic notations, reducing noise and efficiently handling all irregularities" [ 30] . They collaborated with historians and students to corroborate the machine learning outputs, calling it a "cooperative research aid" showing how machine learning research can include humans in the training process. They also created an open-source interface: ithaca.deepmind.com 580 581 582 SOUND ARTIST EXPERIMENTS WITH REFUGEE ACCENT DETECTION TOOLS Pedro Oliveira's work [31] explores the accent recognition software used since 2017 by the German Federal Office for Migration and Refugees (BAMF). Though BAMF does not disclose the software's datasets, Oliveira traced the probable source to two annotated sound databases from the University of Pennsylvania -- unscripted Arabic telephone conversations named "CALL FRIEND" [32] and "CALL HOME" [33] . In 2019 the software had an error rate of 20 percent despite its deployment 9,883 times in asylum 583 seekers' cases [34] . Oliveira utilizes sounds removed from the datasets and reverse engineers the algorithm (as musical transformations rather than for classification tasks), in order to show how politically charged it is to define and detect accents. "How can you say it's an accurate depiction of an accent?" he says. "Arabic is such a mutating language. That's the beauty of it actually" [35] . He presents this through live performance and the online sound essay "On the Apparently Meaningless Texture of Noise" [36] . 584 585 586 HUMAN RIGHTS ACTIVISTS ACCOUNT FOR WAR CRIMES WITH SYNTHETIC DATA Sometimes important training data is missing from a dataset, because not enough of it exists, and these absences can amplify narrow assumptions about a diverse community. In other cases that don't involve human subjects, synthetic data can fill gaps in creative ways. When human rights activists from Mnemonic, who were investigating Syrian war crimes using machine learning, struggled to find enough images of cluster munitions to train 587 their model, computer vision group VFRAME created synthetic data -- 10,000 computer-generated 3D images of the specialized weapon and its blast sites -- which researchers then used to sift through the Syrian Archive's 350,000 hours of video, searching for evidence of war crimes [37] , [38] . Such systems can reduce the number of videos people need to comb through manually, while still keeping humans involved with pattern review and confirmation. 588 THERE ARE MANY MORE EXAMPLES LIKE THESE OF HOW TO SOURCE, USE, 589 AND COMBINE DATASETS THAT ALREADY EXIST. THE CRITICAL AND CREATIVE POSSIBILITIES ARE NEARLY ENDLESS. 590 591 592 593 For your consideration: What kind of dataset(s) will you 594 use, and how can you approach it more critically? How 595 will you apply what you've learned here to your next 596 machine learning project? 597 598 599 600 601 5 602 TRANSFORMING DATASETS 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 Just as there is no such thing as neutral data, no dataset is ready to use off the shelf. From preprocessing (sometimes confusingly called 'cleaning') to model creation, transformations reflect the perspectives of the dataset creators and users. This overview covers some of the technical details of getting a dataset ready for your tasks [18], [23], [26], [39], 623 [40], [41], {9}, while asking critical questions along the way. As artist and researcher Kit Kuksenok argues, "Data cleaning is unavoidable! Each round of repeated care of data and code is an opportunity to invite new perspectives to code/data technical objects" [42]. Preprocessing is a key part of building any system with a dataset, so it is crucial to document and reflect upon preprocessing transformations. 624 625 626 627 CAUTION: Be on the lookout for dataset transformations 628 that result in lost meanings, new misconceptions, or 629 skewed information. 630 631 632 633 634 635 STORING DATA 636 637 638 639 640 A dataset must live somewhere, and once it grows beyond 641 a single manageable file, it usually lives in a 642 DATABASE. While 'dataset' describes what the data are, 643 'database' describes how data are stored -- whether as a 644 set of tables with columns and rows (e.g. a relational 645 database like `SQL`, or a collection of documents with 646 keys and values like `MongoDB`). Database structures 647 should suit what they hold, but they will also shape 648 what they hold and reflect how database designers see 649 data. "They also contain the legacies of the world in 650 which they were designed," says media studies scholar 651 Tara McPherson [43]. 652 653 654 655 656 657 658 659 660 ACCOUNTING FOR MISSING DATA 661 662 663 664 665 666 667 668 You may have entries in your dataset that read `NaN` 669 (not a number) or `NULL`, which may or may not cause 670 errors, depending on what kinds of calculations you do. 671 You may also have manual entries like, '?', 'huh', or 672 blanks that lack context. Should you remove the missing 673 information? If you replace it, how will you know what 674 goes in its place? Is data missing in uniform ways such 675 that whole categories can be eliminated, or is it only 676 missing for subgroups in ways that could skew results? 677 How will you know what impacts your edits may have? 678 Consider what missing data might mean. "Unavailable [is] 679 semantically different from data that was simply never 680 collected in the first place," says data scientist David 681 Mertz [41]. Filtering out data and filling in data have 682 very different implications. Could you consult data 683 subjects to get more context on missing data or the 684 implications of removal or substitutions? How have 685 others handled similar challenges? Can you run tests 686 that treat missing data differently and compare the 687 results? Mimi Onuoha's "The Library of Missing Datasets" 688 reflects on how missing data imply what will not or 689 cannot be collected, or what has been considered not 690 worthy of collection. The project creates a physical 691 archive of empty files, covering topics that are 692 excluded despite our data-hungry culture. She says, 693 "That which we ignore reveals more than what we give our 694 attention to" [44]. 695 696 697 698 699 700 701 702 703 704 705 706 707 HANDLING EXTRA DATA 708 709 710 711 712 713 You'll probably encounter dataset anomalies, outliers, 714 and duplicates and then need ways to identify, adjust, 715 or remove them. In text datasets for unsupervised 716 learning, you'll likely remove punctuation and "stop 717 words" (commonly used conjunctions or articles like 'an' 718 or 'the', for example). But, as software engineer 719 Francois Chollet says, "even perfectly clean and neatly 720 labeled data can be noisy when the problem involves 721 uncertainty and ambiguity" [18]. Outliers can also be 722 accurate and contain meaningful information. As Crawford 723 emphasizes, such acts of data cleaning and 724 categorization create their own concepts of outside and 725 otherness that can restrict "how people are understood 726 and can represent themselves" [25]. Defining outliers, 727 anomalies, or extra data means deciding what is 728 'normal', unexpected, or distracting -- what is signal 729 and what is noise. 730 731 732 733 734 735 736 737 738 739 DISCRETIZING DATA: 740 741 742 743 744 AKA "binning" or grouping instances together may be 745 useful when you don't need the original level of detail 746 provided (see 'dimensionality reduction' in ENGINEERING 747 FEATURES, below). For example, you might switch 748 continuous data like temperature readings into bins 749 grouped by every five or ten degrees. There are built-in 750 functions for doing so, but remember that creating data 751 ranges can skew results and may not be appropriate for 752 all cases. Make sure to document any changes and provide 753 the original dataset as well as the modified version. 754 755 756 757 758 759 760 761 762 763 TOKENIZING OR CHUNKING DATA: 764 765 766 767 768 Breaking up data into smaller units. Tokens are often 769 individual words or sentences. Other text-related 770 operations might include removing punctuation and stop 771 words that are commonly used (see HANDLING EXTRA DATA). 772 Datasets should use tokenization strategies that account 773 for linguistic differences, since a 'word' as a unit of 774 meaning can vary significantly among languages. Other 775 sequence-like data types like audio and video are also 776 broken up into chunks, sometimes in order to be 777 processed, compressed, or streamed. 778 779 780 781 782 783 784 785 786 NORMALIZING DATA 787 788 789 790 791 Altering numerical data to bring them all within the 792 same range and using the same units, e.g. between 0-1, 793 is called normalization, or scaling [18]. In text 794 datasets, this can also mean converting all text to 795 lowercase and standardizing each word by reducing it to 796 its root form (stemming or lemmatizing). In image data, 797 this could also mean cropping images to the same 798 dimensions or around the same subject, changing the 799 color profile to grayscale, and so on. Remember that, 800 while useful for many tasks, normalizing data 801 potentially removes important context or adds ambiguity 802 to the data (e.g. cropped images ignore any information 803 outside the frame, acronyms may be read as their 804 homonyms, scaled numbers may then be rounded and lose 805 specificity). 806 807 808 809 810 811 812 813 814 815 SEPARATING TESTING & VALIDATION DATA 816 817 818 819 820 If it has not already been separated, mark off a portion 821 of your data that will not be exposed to your model or 822 used to train it in any way. Keep these testing and 823 validation portions separate from your training data, so 824 that they can be used later to test your model's 825 performance on new information once it has been trained 826 on the remaining training data. (See Section 3, TRAINING 827 DATA and VALIDATION DATA for more context.) 828 829 830 831 832 833 834 835 836 ENGINEERING FEATURES 837 838 839 840 841 842 843 You may need to create features (e.g., add columns to 844 your table) to show data from new perspectives. This can 845 impact how the dataset can be analyzed going forward, 846 how the model can be designed, and how the data subjects 847 and subjectees might be affected. For example, the 848 unsupervised machine learning approach called 849 'dimensionality reduction' analyzes a dataset for its 850 most relevant features so that the rest can be ignored. 851 It can also involve combining or altering existing 852 features to simplify the whole. However, this runs 853 directly counter to what legal scholar Kimberle Crenshaw 854 calls intersectional analysis -- an approach that rejects 855 grouping people together in categories without attending 856 to the unique experiences (and the data) of people for 857 whom those categories intersect, who are most negatively 858 impacted by systems with the power to categorize people 859 [45] (For more on intersectionality, see ALTERNATIVE 860 APPROACHES TO DATASET PRACTICES). 861 862 863 864 865 866 867 868 869 870 871 EXPLORING DATA 872 873 874 875 Sorting, sampling, combining, pivoting, and visualizing 876 data are other transformations you will likely use 877 during dataset preprocessing. These will differ greatly 878 depending on the type of dataset and your project's 879 objectives, but they all require asking critical 880 questions about how such explorations influence the 881 meaning of the dataset, the model developed, and the 882 system deployed. 883 884 885 886 887 888 889 890 891 892 For your consideration: How have your data 893 transformations shaped your dataset so far? Which 894 transformations are most thought-provoking and worth 895 exploring? Who else could offer perspective on your 896 preprocessing decisions? 897 898 899 900 901 902 6 903 THE DATASET'S LIFECYCLE 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 DATASETS OFTEN HAVE SURPRISING HISTORIES, USES, AND AFTERLIVES. FROM LOCATING THE BEST DATASET FOR THE JOB, TO WORKING WITH 921 IMPERFECT EXISTING DATA, TO SHARING RESULTS AND ACCOUNTING FOR IMPACTS, HERE ARE SOME CRITICAL QUESTIONS TO ASK AT EACH STAGE OF A PROJECT THAT USES MACHINE LEARNING DATASETS: 922 923 924 THE DATASET'S LIFECYCLE 925 6.1 926 ORIGINS: WHAT IS YOUR DATASET'S STORY? 927 928 929 "Datasets are the results of their means of collection," says artist and technology researcher Mimi Onuoha [46]. They are influenced by the many people who contribute to them, who participate in their creation (knowingly or unknowingly), who collect data, who annotate or label data, or who are affected by a machine learning system that uses the dataset. 930 These questions will help you select a dataset to work with, and to understand how its creation could inform your project if you use it. Often the answers to these questions can be found in a dataset's DATASHEET or a related research paper. Although including datasheets is becoming a standardized practice, not all datasets have datasheets, and many datasheets are incomplete. Look for datasets with complete datasheets, updated documentation, and current contact information for its creators. 931 932 933 934 WHO CREATED THIS DATASET? WHO FUNDED IT? WHAT WERE THEIR 935 MOTIVATIONS OR AIMS, AND HOW DO THEY COMPARE TO YOURS? 936 [24] 937 938 939 940 941 942 943 If they differ in significant ways, consider how using 944 the dataset for other purposes will impact both the 945 original data subjects, as well as data subjectees, and 946 the outcomes of your project. Would an alternative 947 dataset be more appropriate? Document the rationale for 948 the dataset you choose. 949 950 951 952 953 954 955 956 957 958 HOW WAS THE DATA COLLECTED? HOW WAS IT ANNOTATED AND BY 959 WHOM? ARE THE ORIGINAL ANNOTATION INSTRUCTIONS 960 AVAILABLE? HOW ARE THE LIMITATIONS OF THOSE METHODS 961 ACCOUNTED FOR? [47] WERE DATA SUBJECTS PART OF THE 962 DATASET'S DESIGN AND CREATION? WAS THE RESULTING DATA 963 VALIDATED BY ITS SUBJECTS? [24], [48] 964 965 966 967 968 969 970 971 972 If information about collection and annotation is 973 missing, or if the collection and annotation methods 974 were inappropriate or misaligned with your objectives, 975 you may want to consider an alternative dataset. Also 976 consider the contexts of labeling and annotation, as 977 crowd-sourced data can lack the nuance of individual 978 annotators from diverse perspectives who had the 979 opportunity to collaborate [49]. 980 981 982 983 984 985 986 987 988 989 HOW HAS THE DATASET BEEN PROCESSED ALREADY? IS IT A 990 SMALL SAMPLE OF A LARGER COLLECTION? IS IT A COMPILATION 991 OF OTHER PRE-EXISTING DATASETS? HAS IT BEEN STANDARDIZED 992 OR TRANSFORMED IN ANY WAY (SEE SECTION 5)? 993 994 995 996 997 998 999 1000 1001 If the dataset is a sample from a larger dataset or a 1002 compilation of smaller datasets, investigate those 1003 original sources to see if its data matches the sample 1004 or is more appropriate for your work. Does information 1005 in its datasheet impact your decision to use this 1006 dataset? If it has been transformed, see if the 1007 documentation also includes the original version or a 1008 description of its methods. 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 WHAT DOES THE DATASET CONTAIN? DOES IT INCLUDE A 1019 CODEBOOK DESCRIBING ITS PARTS? WHICH PERSPECTIVES ARE 1020 INCLUDED, AND WHICH ARE MISSING? WHICH OUTLIERS ARE 1021 DISMISSED, AND WHAT DATA IS UNACCOUNTED FOR? CAN YOU 1022 AUDIT THE DATASET, OR HAS IT ALREADY BEEN AUDITED? WHAT 1023 DOES THE AUDIT SHOW AND HOW CAN YOU ACCOUNT FOR ITS 1024 FINDINGS? 1025 1026 1027 1028 1029 1030 1031 1032 If the dataset has gaps that neglect important 1033 considerations or that might affect your project, would 1034 a different dataset be more appropriate? 1035 1036 1037 1038 1039 1040 Compare with other datasets in this area to see how they 1041 account for similar issues. Consult with your project 1042 group or community of practice to see how they interpret 1043 these issues and their importance. Get their help to 1044 spot issues you might have missed, and consider any 1045 other datasets that might better fit the project. Offer 1046 in-kind support. 1047 1048 1049 1050 1051 1052 1053 Regardless of whether you use the dataset, be sure to 1054 document your questions and concerns. Describe the 1055 limitations you see in the dataset, discuss their 1056 relevance to your project, and share what you have done 1057 to mitigate their impacts. Proceed with caution, if at 1058 all. 1059 1060 1061 1062 1063 1064 1065 1066 1067 WHEN WAS THE DATASET MADE? IS THIS ITS LATEST VERSION? 1068 IF IT HAS BEEN DEPRECATED (OR REMOVED FROM PUBLIC 1069 CIRCULATION), WHY? DOES IT CONTAIN INFORMATION THAT IS 1070 INACCURATE OR OFFENSIVE? FOR MORE DISCUSSION OF 1071 DEPRECATED DATASETS, SEE SECTION 7.2.3. 1072 1073 1074 1075 1076 1077 1078 1079 If the dataset is no longer valid, you will need to find 1080 another dataset. Maybe an updated version of the dataset 1081 addresses the issues that led to its removal, or perhaps 1082 you can make these revisions yourself. However, 1083 resolving ethics or accuracy issues is not as simple as 1084 updating a table, since the underlying structure of the 1085 dataset may remain problematic. Proceed with caution, if 1086 at all. 1087 1088 1089 1090 1091 Just because a dataset is *not* deprecated does not mean 1092 it is safe to use. There is not yet any standardized way 1093 to audit, update, or remove existing datasets, or even a 1094 central repository where issues can be documented and 1095 addressed. Much remains at the creators' discretion [50] 1096 . 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 WHO IS FEATURED IN THE DATASET? WHO IS LEFT OUT? HOW 1108 DOES THE DATASET ACCOUNT FOR WHAT IS MISSING? WHAT 1109 ASSUMPTIONS, INTUITIONS, THEORIES, STEREOTYPES, OR 1110 INEQUITIES ARE CONTAINED IN THE DATA OR BUILT INTO THE 1111 DATASET'S STRUCTURE? HOW MIGHT THESE FRAMEWORKS HAVE 1112 BEEN PERPETUATED THROUGH FORMATTING AND TRANSFORMATION 1113 PROCESSES AS THE DATASET WAS MADE MACHINE-LEGIBLE? 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 If you are unsure how the dataset and your use of it may 1124 impact a community, consult with its members to 1125 understand their perspectives. "Build with, not for," 1126 says the Design Justice Network, which emphasizes that 1127 "community members already know what they need and are 1128 working towards solutions that work for them" [51]. 1129 1130 1131 1132 1133 Consider the consequences of inclusion. In an unjust 1134 system (e.g. racist criminal sentencing) more 1135 representation is not the right answer. Completeness may 1136 do more harm. 1137 1138 1139 1140 Document your processes. Noticing what information is 1141 left out can be just as important as what is shown [48]. 1142 If too much information is missing, or omissions are too 1143 problematic, you may need to find another dataset. 1144 1145 1146 1147 1148 Maybe just don't build it. The best solution to a 1149 problem is not always more technology or a new system. 1150 Explaining why you did not build a model or did not use 1151 a dataset can still be a valuable contribution - showing 1152 the need for caution, skepticism, and alternative 1153 approaches. 1154 1155 1156 1157 1158 1159 1160 1161 HOW WAS CONSENT GIVEN FOR INCLUSION IN THE DATASET? WAS 1162 CONSENT FULLY INFORMED, VOLUNTARY, AND REVOCABLE? HOW 1163 ARE SUBJECTS' ANONYMITY PROTECTED? 1164 1165 1166 1167 1168 1169 1170 If the dataset does not adequately address how consent 1171 was provided, or if this consent does not extend to the 1172 uses of your project, you will need to find a different 1173 dataset. If your use potentially risks subjects' 1174 anonymity as established in the original dataset design, 1175 you will need to find a different dataset. Privacy 1176 cannot be guaranteed; be wary of the potential for 1177 re-identification of so-called anonymous data when they 1178 are combined with other datasets [52]. 1179 1180 1181 1182 1183 1184 If full consent was not given -- especially if consent 1185 was not possible to obtain -- "just don't build it" is 1186 always a valid and responsible option. 1187 1188 1189 1190 1191 1192 1193 DO YOU HAVE LEGAL, ETHICAL ACCESS TO THIS DATASET? DOES 1194 YOUR USE OF IT ALIGN WITH ITS LICENSING AND TERMS OF 1195 USE, AND WITH YOUR OWN CODES OF CONDUCT? ARE ANY OF THE 1196 DATA INTENDED FOR RESTRICTED USE BY SPECIFIC 1197 COMMUNITIES? [53] - [55] 1198 1199 1200 1201 1202 1203 1204 1205 We are not lawyers and this is not legal advice, of 1206 course. Check at the beginning, middle, and toward the 1207 end of your project -- before publication or launch -- 1208 that your use is authorized and appropriate to its 1209 creators and its subjects. It should be in keeping with 1210 both the letter and the spirit of terms of use and codes 1211 of conduct. Talk with others doing similar work to see 1212 what pitfalls they wish they had avoided upfront. 1213 1214 1215 1216 1217 OVERALL, IS THIS THE BEST DATASET FOR YOUR PROJECT? WHAT ARE THE TRADEOFFS OF USING THIS DATASET VERSUS ANOTHER ONE? 1218 1219 1220 THE DATASET'S LIFECYCLE 1221 6.2 1222 USAGE: WHAT IS THE STORY YOU WILL TELL WITH YOUR DATASET? 1223 1224 1225 Your project aims will steer your dataset selection, the features you choose to interpret, the model you pick, and the 1226 adjustments you make. Across all of these small and large technical choices, you can use critical lenses to achieve your aims and minimize harms. 1227 1228 1229 WHAT IS YOUR PROJECT'S GOAL? HOW DOES THE DATASET HELP 1230 YOU ACHIEVE IT? 1231 1232 1233 1234 Prioritize your own purpose and approach over the 1235 popularity of the dataset or your familiarity with it. 1236 What is the best dataset for this task, to answer these 1237 questions? 1238 1239 1240 1241 1242 1243 1244 1245 WHAT ASSUMPTIONS ARE YOU PRIORITIZING OR EXCLUDING BY 1246 HOW YOU'VE FRAMED YOUR PROJECT? HOW MIGHT IT BE REFRAMED 1247 TO GAIN MORE INSIGHT? HOW MIGHT YOU COLLABORATE WITH 1248 PEOPLE FROM DIFFERENT DISCIPLINES OR BACKGROUNDS -- E.G. 1249 ARTISTS, PRACTITIONERS, OR COMMUNITY STAKEHOLDERS - TO 1250 DEVELOP THE PROJECT'S AIMS MORE RICHLY? 1251 1252 1253 1254 1255 1256 1257 1258 Gather and prioritize the perspectives of data subjects 1259 and data subjectees. Connect with experts from other 1260 domains for fresh eyes and constructive feedback. Seek 1261 out information that you would not normally encounter; 1262 don't automatically rule out information in an 1263 unfamiliar vocabulary or format. 1264 1265 1266 1267 1268 1269 1270 1271 WHAT ASSUMPTIONS ARE YOU MAKING AS YOU PROCESS AND CLEAN 1272 THE DATASET, AND AS YOU SELECT FEATURES FOR ANALYSIS? 1273 COULD CATEGORIES (FEATURES) IN THIS DATASET BE 1274 DISAGGREGATED TO TELL A DIFFERENT STORY THROUGH THE 1275 DATA? SOME COMMON ASSUMPTIONS: 1276 1277 1278 1279 1280 1281 How have you treated the collected data as neutral in 1282 any obvious or more subtle ways? 1283 1284 1285 1286 Consider how transforming data [40] during the cleaning 1287 process may misrepresent information or remove important 1288 detail from the dataset. e.g. `NaN` (Not a Number) may 1289 conceal data that was never collected. (See Section 5 1290 for more.) 1291 1292 1293 1294 A dataset is frequently treated as a generalizable, 1295 multi-purpose tool, applicable across many tasks and 1296 disciplines, when more often it is best suited or only 1297 suited to its original purpose. 1298 1299 1300 Data cleaning is not one-and-done, but an iterative, 1301 integral process [42]. Though often undervalued labor, 1302 data cleaning is part of the 'real' work of making 1303 datasets. 1304 1305 1306 1307 1308 1309 WHAT ASPECTS OF THE DATASET WILL YOU INCLUDE OR EXCLUDE 1310 IN YOUR PROJECT, AND HOW DO THEY CONVEY THE INFORMATION? 1311 HOW WILL YOU ENSURE THESE CHOICES DO NOT OVERLY SKEW THE 1312 RESULTS? 1313 1314 1315 1316 1317 1318 FEATURE SELECTION & ENGINEERING. Distilling a dataset 1319 into pertinent columns is an essential part of dataset 1320 work because it determines what information categories 1321 will be important for later analysis. This process is 1322 descriptive and creative, not self-evident [26]. 1323 (See Section 3: FEATURE and Section 5: ENGINEERING 1324 FEATURES for more.) 1325 1326 1327 1328 1329 1330 HOW FAR DOES YOUR PROJECT DEVIATE FROM THE DATASET'S 1331 ORIGINAL PURPOSE? IF IT IS SIGNIFICANTLY DIFFERENT, HOW 1332 WILL YOU TRACK ANY NEW IMPACTS? [54] 1333 1334 1335 1336 1337 1338 1339 When diverging from the original objectives of a 1340 dataset, ensure that your dataset transformation 1341 processes align with both your own goals and any 1342 guidelines put in place by the dataset's creators. 1343 Consider consent and licensing restrictions, as well as 1344 other potential legal and ethical issues. Furthermore, 1345 what kinds of impacts would the creators not have 1346 foreseen in your use of their dataset? How might 1347 communities be newly impacted by your use of this 1348 dataset, and how can you engage them in this process? 1349 1350 1351 1352 1353 1354 1355 COULD YOUR USE OF THE DATASET CAUSE HARM? HOW WILL YOU 1356 MEASURE AND ADDRESS ADVERSE IMPACTS? WHAT STEPS WILL YOU 1357 TAKE TO MINIMIZE HARM? [54] 1358 1359 1360 1361 1362 Work closely with data subjects and potential data 1363 subjectees who may be impacted by your use of the 1364 dataset, in order to discover what mitigation strategies 1365 would be best for them. How you address potential 1366 impacts will depend greatly on the types of risks your 1367 project presents; listening across diverse communities 1368 and disciplines will help you uncover new issues, create 1369 checkpoints, and take useful action. Encourage in your 1370 project team a culture of openness and learning from 1371 mistakes [48]. 1372 1373 1374 1375 1376 1377 1378 HOW WILL YOU MAINTAIN THE CONSENT AND ANONYMITY OF ANY 1379 DATA SUBJECTS? HAVE THEY BEEN TOLD ABOUT THE RISKS 1380 SPECIFIC TO YOUR USE CASE? CAN THEY CHECK BACK ON HOW 1381 THEIR DATA HAS BEEN USED? 1382 1383 1384 1385 1386 1387 Consent requires that subjects understand both the 1388 implications of data's use and impact, as well as their 1389 digital rights [48]. This goes beyond terms of service 1390 disclosures and should include the purpose of the 1391 project, any risks, and instructions on how to revoke 1392 consent if necessary. 1393 1394 1395 1396 There is no such thing as "anonymizing" identifying data 1397 because data can easily be combined with other sources 1398 to re-identify people [48]. 1399 1400 1401 1402 1403 1404 1405 DOES YOUR WORK WITH THIS DATASET RESULT IN A NEW, 1406 DERIVATIVE DATASET? HOW WILL YOU ACCOUNT FOR NEW ETHICAL 1407 CONCERNS ARISING FROM THE DERIVATIVE DATASET WHILE STILL 1408 ADDRESSING ISSUES RAISED BY THE ORIGINAL? 1409 1410 1411 1412 1413 1414 Refer to Section 6.1 ORIGINS, as well as to resources 1415 for creating datasets like Gebru, et al.'s, "Datasheets 1416 for Datasets" to ensure that you are considering the 1417 questions that arise from creating a new or derivative 1418 dataset [24]. Your documentation, including a new 1419 datasheet, will help others who may want to use your new 1420 dataset. 1421 1422 1423 1424 1425 OVERALL, HOW HAS YOUR INITIAL EXPLORATION OF THE DATASET CHANGED 1426 YOUR PROJECT OR ITS GOALS? WHAT NEW QUESTIONS DOES IT RAISE? 1427 1428 1429 1430 THE DATASET'S LIFECYCLE 1431 6.3 1432 STEWARDSHIP: WHAT STORY WILL THIS DATASET KEEP TELLING? 1433 1434 1435 Although you may have completed your analysis, your work with the dataset is not done. Critical stewardship takes a holistic 1436 approach to sharing, maintaining, and deprecating a project. It's not just fixing something when it breaks, it requires sustainable and thoughtful relationships. Dataset stewardship lasts the whole data lifecycle. 1437 1438 1439 1440 1441 HOW WILL YOU SHARE THIS DATASET? WILL YOU PROVIDE ACCESS 1442 TO YOUR MODIFIED VERSION OR LINK TO THE CREATORS' 1443 ORIGINAL VERSION? WHO WILL HAVE WHAT KINDS OF ACCESS 1444 (OPEN-SOURCE, AUTHENTICATED, PLATFORM-BASED)? [54] 1445 1446 1447 1448 1449 1450 1451 1452 1453 Maintaining any existing terms of service or consent 1454 agreements, consider making your dataset and results as 1455 available as possible. Keep in mind that the spirit of 1456 open access means more than just uploading your files to 1457 a repository or posting a link, it also includes sharing 1458 clear, complete documentation. Consider audiences 1459 outside your project team, including the data subjectees 1460 (see Section 3) and others who may be impacted, and 1461 include instructions for how to use the dataset in plain 1462 language. Code notebooks, examples, screenshots, and use 1463 cases are all helpful; and you can support broader 1464 access by using open-source, free tools, and multiple 1465 formats for creating your examples. 1466 1467 1468 1469 1470 1471 1472 1473 ARE YOUR RESULTS FAIR (FINDABLE, ACCESSIBLE, 1474 INTEROPERABLE, REUSABLE)? [56] 1475 1476 1477 1478 1479 List your project findings and dataset with dataset 1480 repositories. Many repositories have a section to list 1481 multiple projects that cite a particular dataset. Make 1482 sure that your dataset has a DOI and complete metadata, 1483 and that files are in standard formats. Include a clear, 1484 comprehensive license so that it can be reused 1485 appropriately. {10} 1486 1487 1488 1489 1490 1491 1492 HOW WILL YOU DOCUMENT ANY ADDITIONAL DATA PREPROCESSING 1493 THAT WAS NECESSARY FOR YOUR USE OF THE DATASET? WHAT 1494 OTHER KINDS OF DOCUMENTATION (DATASHEETS, CODEBOOKS, 1495 ETC.) ARE NECESSARY? [24] 1496 1497 1498 1499 1500 1501 1502 Because you are using your dataset for a new project 1503 with a new objective, it makes sense to create a new 1504 datasheet. Cite the original datasheet and make any 1505 adjustments that reflect your project. E.g. if you 1506 created a new feature to study whether or not a user was 1507 answering a survey online, and this added a column to 1508 your version of the dataset, include that in the 1509 datasheet. Discuss why each decision was made and how 1510 the work was done. It may seem tedious, but keeping a 1511 research notebook or a working draft of your datasheet 1512 as you go can become a regular part of your practice, an 1513 easy way to document your work, and great help to other 1514 people who work with your dataset in the future. 1515 1516 1517 1518 1519 1520 1521 1522 WHAT PRESENTATION FORMS WILL YOU USE TO TELL STORIES 1523 WITH THIS DATASET? CAN YOU COLLABORATE WITH OTHERS TO 1524 USE DIFFERENT FORMS OR ENGAGE DIFFERENT SENSES? [2] 1525 1526 1527 1528 1529 1530 So much meaning is encoded in seemingly simple design 1531 choices. Make those choices intentionally, work with 1532 people who specialize in data-based storytelling, and 1533 value input and opportunities to reach new audiences 1534 with different forms. The Design Justice Network has a 1535 zine collection illustrating how to practice 1536 community-centered, equitable design [51]. 1537 1538 1539 1540 1541 1542 1543 HOW WILL YOU MAINTAIN AND MONITOR ACCESS TO YOUR DATASET 1544 IN WAYS THAT CONSIDER THE INTELLECTUAL PROPERTY AND 1545 PRIVACY RIGHTS OF DATA SUBJECTS AND SUBJECTEES (SEE 1546 SECTION 3)? 1547 1548 1549 1550 1551 1552 1553 If your dataset has privacy or consent constraints, 1554 cultural considerations, proprietary constraints, or 1555 other reasons that it should not be shared broadly, make 1556 a plan for data storage. This plan should be as secure 1557 (if not more) and long-standing as the plan for the 1558 original dataset. Create a stewardship chain for the 1559 project and its infrastructure that will be maintained 1560 after you or your team have moved on [50]. 1561 1562 1563 1564 1565 1566 1567 HOW WILL YOU KNOW IF THE DATASET'S CREATORS REVISE OR 1568 DEPRECATE THE DATASET, AND WHAT IS YOUR PLAN FOR 1569 HANDLING SUCH CHANGES? [50] 1570 1571 1572 1573 1574 In your data stewardship plan, include regular checks of 1575 the original dataset's website or repository. Know how 1576 you will proceed if the dataset is revised or removed. 1577 This may mean revising your own dataset, revising your 1578 preprocessing, updating models, or reconfirming impacts 1579 to community members and their consent. 1580 1581 1582 1583 1584 1585 1586 WHO WILL ARCHIVE AND/OR DEPRECATE YOUR OWN DATA WHEN 1587 NECESSARY, AND HOW WILL THIS BE DONE? 1588 1589 1590 1591 1592 1593 1594 Follow deprecation best practices using guidelines like 1595 the ones recommended by Luccioni and Corry, et al.[50] 1596 (discussed in Section 7). Their "Framework for 1597 Deprecating Datasets" suggests including the reasons for 1598 deprecation, how the removal will occur and plans for 1599 mitigating any negative impacts, an appeal mechanism for 1600 others who may be making use of your work, a timeline of 1601 the process, protocols for access after deprecation 1602 (usually with restrictions for research, legal, or 1603 historical use only), and a publication check request 1604 asking future paper authors to confirm they are not 1605 using a deprecated version. 1606 1607 1608 1609 1610 1611 1612 1613 1614 For your consideration: Although principles like 1615 indigenous data governance or data feminism may seem 1616 abstract or hard to apply, they illustrate practices 1617 that can help you design better and more thoughtful 1618 projects that accomplish your goals and respect people 1619 who have historically been excluded from and harmed by 1620 dataset design. 1621 1622 1623 1624 HOW DOES YOUR WORK ALIGN WITH PRINCIPLES OF DATA FEMINISM - 1625 I.E., EXAMINING AND CHALLENGING POWER, ELEVATING EMOTION AND EMBODIMENT, RETHINKING BINARIES AND HIERARCHIES, EMBRACING PLURALISM, CONSIDERING CONTEXT, AND VALUING LABOR? [47] 1626 ARTICULATED IN THE CARE PRINCIPLES FOR INDIGENOUS DATA GOVERNANCE, HOW DOES THE DATASET OFFER COLLECTIVE BENEFIT; GIVE 1627 ITS SUBJECTS AUTHORITY TO CONTROL DATA; REQUIRE RESPONSIBILITY FROM PROJECT TEAMS; AND CENTER ETHICS, HUMAN RIGHTS, AND WELLBEING "AT ALL STAGES OF THE DATA LIFE CYCLE AND ACROSS THE DATA ECOSYSTEM"? [57] 1628 1629 1630 7 1631 CAUTIONS & REFLECTIONS FROM THE FIELD 1632 1633 1634 1635 1636 1637 1638 1639 1640 1641 1642 1643 1644 1645 1646 1647 1648 1649 How does mishandling datasets contribute to harm? Like any messy, multifaceted material, datasets must be treated with care. Taking time to see the broader implications of making and using datasets can save you time, create projects that are 1650 easier to explain, and help you build stronger relationships with the communities your datasets impact. Here we review why datasets matter to machine learning, why current approaches can sometimes be inadequate, and some lessons learned from those who have worked with machine learning datasets. 1651 1652 1653 CAUTIONS & REFLECTIONS FROM THE FIELD 1654 7.1 1655 DATASETS DIRECTLY IMPACT LIVES 1656 1657 1658 Improper dataset use can impact DATA SUBJECTS (people contained in the original dataset), DATA SUBJECTEES (people analyzed or affected by a dataset's machine learning system, see Section 3), DATA WORKERS (the laborers who prepared the dataset), and even 1659 DATA RESEARCHERS, DESIGNERS, JOURNALISTS, ENGINEERS, ARTISTS, or other communities working with datasets. Whether the machine learning systems driven by datasets help deny resources to individuals (allocative harms), or misrepresent communities (representational harms) [58], they can powerfully and dangerously impact people. 1660 Algorithmic processes often mirror and magnify existing power imbalances in social systems. In a process that sociologist Ruha Benjamin calls "coded inequity," algorithms and datasets "reflect and reproduce existing inequities" while simultaneously promoting them as "more objective or progressive." The speed and 1661 scale of machine learning and massive datasets make "discrimination easier, faster, and even harder to challenge" [59]. It also makes inequity more invisible and insidious, and dataset users can best understand the implications of their materials by looking carefully for impacts and working with others who can see impacts they might miss. 1662 1663 1664 CAUTIONS & REFLECTIONS FROM THE FIELD 1665 7.2 1666 WHERE DO DATASETS GO WRONG? 1667 1668 1669 Whether designing a dataset from scratch or using one that has been around for years, decisions made at every step will inform your project outcomes. These decisions get scaled and compounded by machine learning models. This section summarizes common 1670 pitfalls of working with existing datasets and suggests ways to be cautious at various parts of a dataset project. In a survey of how machine learning researchers often work with large, complex datasets, natural language processing researcher Amandalynne Paullada, et al., found four kinds of COMMON DATASET PITFALLS (see box). 1671 1672 1673 COMMON DATASET PITFALLS 1674 1675 spurious tasks: "where success is only possible [...] 1676 because the tasks themselves don't correspond to reasonable real-world correlations or capabilities" 1677 1678 artifacts in the data: "which machine learning models can easily leverage to 'game' the tasks" 1679 sloppy annotation or documentation: a lack of reflective 1680 description can "erode the foundations of any scientific inquiry based on these datasets" 1681 representation: "wherein datasets are biased both in 1682 terms of which data subjects are predominantly included and whose gaze is represented" [60], {11} 1683 1684 Even those actively trying to fix datasets can experience these pitfalls. Paullada, et al. observed that datasets which were modified after their creation -- often in attempts to improve a model's ability to generalize -- were still susceptible to the same kinds of problems as the originals. They suggest "a broader 1685 view to be taken with respect to rethinking how we construct datasets for tasks overall," including dataset cultures around benchmarking, use and reuse, and licensing [60]. Importantly, they emphasize the need to move beyond technical fixes to consider dataset stewardship holistically, from project design to deprecation, from historical and technical foundations to field-wide approaches. 1686 During dataset creation, classification thinking already shapes how data is collected, organized, and later perceived. In a multi-year analysis of outsourced data work, computer and information science researchers Milagros Miceli and Julian Posada found that the majority of classification choices for crowdworkers who are labeling data were simple 'Either-Or' selections, with little room for complexity. Workers were "encouraged to ignore ambiguity altogether" [62]. {12} While binary thinking simplifies data collection and labeling, it 1687 codifies the viewpoints of those who create the binaries: "These examples of social classification and conceptualization are not just about cultural differences between requesters and data workers, but they reflect the prevalence of worldviews dictated by requesters and considered self-evident to them" [62]. It is critical to remember that crowdworkers' judgments and abilities to discern complexity are silenced in classification tasks that leave no room for debate or discussion, making the resulting classifications "cleaner" but far less valuable for appreciating the richness of data. 1688 After datasets are created, they are often not audited, revised, or maintained. In many cases, they continue to be used despite having serious flaws. When datasets are reassessed, they might be deprecated for legal, organizational, social, or technical reasons. However, as machine learning researcher Sasha Luccioni and critical data scholar Frances Corry, et al., found (as part of the Knowing Machines project, from which this Guide also comes), frequently the processes of deprecation lacked the communication and transparency needed to dissuade usage effectively. They traced six popular datasets that continued to be used after they had been deprecated for privacy violations, 1689 offensive language and imagery, problematic descriptive categories, ethics board violations, and lack of consent uncovered by investigative journalism [50]. Sometimes dataset creators simply move on from their roles or the infrastructure does not exist to sustain the dataset. However, these "zombie datasets" continue to create problems when they keep circulating and feeding machine learning models. While continued use of datasets as training data for machine learning models is one danger, datasets may also persist when incorporated into other datasets [50]. Just because a dataset is available, do not assume that it is without problems, or that it has not been deprecated. 1690 Datasets need ongoing stewardship and reassessment. They must be re-evaluated over time, due to changes in laws and across different jurisdictions [50]. Anyone who uses such datasets, even without knowing their problems, could be at risk, say 1691 Luccioni and Corry, et al., citing Google, Microsoft, and Amazon's legal repercussions for use of IBM's "Diversity in Faces" dataset [50]. Datasets also change context over time due to cultural changes (sometimes called "semantic drift") or repurposed uses that make their data irrelevant, inappropriate, or harmful. {13} 1692 1693 1694 CAUTIONS & REFLECTIONS FROM THE FIELD 1695 7.3 1696 WHY NOT SIMPLY 'DE-BIAS' DATASETS? BECAUSE "BIAS" IS ALWAYS BUILT-IN 1697 1698 1699 Harms cannot be eliminated completely by removing problematic content, optimizing the system, or finding the perfect dataset. Although definitions of BIAS (both technical definitions and its varied cultural understandings) drive many conversations about accountability, diversity, fairness, ethics, explainability, and transparency in machine learning, "bias" is a complicated term. In a survey of almost 150 papers trying to address bias in machine learning, computer scientist Su Lin Blodgett, et al., 1700 found that authors struggled to reach consensus on definitions of "bias" and to articulate how the biased systems were harmful and to whom [66]. The term stands in for a complex set of concerns embedded more foundationally in machine learning systems, often centering on classification. Much interdisciplinary research examines classification -- both its history and function as a fundamental mechanism of machine learning and also how its core principles operate technologically and socially -- which this guide cannot fully address here.{14} 1701 While important research in machine learning techniques is investigating how to create more robust models - whether by refining learning techniques, supplementing data with different or adversarial datasets, or applying other approaches [72] - 1702 [74], {15}- these quantitative strategies can patch key issues or improve accuracy, but they cannot fix underlying structural issues or embedded sociotechnical problems [75], [76]. Rather than attempting to remove bias or avoid classification altogether, work to move beyond bias-focused quick fixes. This guide recommends striving for a layered approach. 1703 1704 1705 AWARENESS: 1706 The people and other beings whose data are gathered into a 1707 dataset. Even if identifying information has been removed, datasets are still connected to the subjects they claim to represent. 1708 1709 1710 ACTION: 1711 Second, be prepared to shift your project if the classifications in your dataset are potentially too reductive, oppressive, or 1712 harmful -- such that the dataset should not be used for machine learning. For situations in which the resulting system might contribute to power imbalances, or collapse complex identities or relations, just don't build it. 1713 1714 1715 1716 For your consideration: While it may be impossible to 1717 escape classification's worldviews entirely, with 1718 awareness of the underlying assumptions of 1719 classification and its impact on your processes, it 1720 becomes easier to make critical decisions that account 1721 for these contexts. 1722 1723 1724 1725 1726 1727 INTERSECTIONAL APPROACHES TO DATASET PRACTICES 1728 1729 1730 1731 1732 1733 1734 Many who work with datasets are already building 1735 alternative systems and strategies. Machine learning can 1736 be approached with fundamentally different mindsets and 1737 aims from the start. One approach being used to address 1738 this question is INTERSECTIONALITY, which is grounded in 1739 Black feminism and the legal theory of Kimberle 1740 Crenshaw. Intersectionality analyzes how power operates 1741 at system-wide scales, sustaining oppression and shaping 1742 identities in layered ways. Intersectionality is also a 1743 set of active strategies developed by communities and 1744 passed down over time [43], [77] - [81]. Intersectional 1745 principles applied to datasets include centering those 1746 who have been at the margins and those impacted most, 1747 maintaining their priority at each phase of a dataset's 1748 lifecycle, and emphasizing ethics of relationality and 1749 care [81] - [83]. (For a variety of perspectives 1750 applying intersectionality to digital technologies, see 1751 the anthology The Intersectional Internet edited by 1752 internet researcher Safiya U. Noble and professor of 1753 education and psychology Brendesha M. Tynes [81] are 1754 some more strategies: 1755 1756 1757 1758 1759 1760 1761 1762 1763 Listen: AI researcher Pratyusha Kalluri argues, 1764 "Researchers should listen to, amplify, cite and 1765 collaborate with communities that have borne the brunt 1766 of surveillance: often women, people who are Black, 1767 Indigenous, LGBT+, poor or disabled." Rather than 1768 focusing on fairness, they ask, "How is AI shifting 1769 power?" [84] 1770 1771 1772 1773 1774 1775 1776 Question assumptions, engage consequences: Writer and 1777 data ethics researcher Anna Lauren Hoffman suggests that 1778 practitioners engage the "consequences of our work, but 1779 also our assumptions, our categories, and our position 1780 relative to the subjects of the data we work with" [85]. 1781 1782 1783 1784 1785 1786 Put power in context: Information science researchers 1787 Miceli, Posada, and Yang recommend contextualized 1788 power-aware approaches that account for "historical 1789 inequities, labor conditions, and epistemological 1790 standpoints inscribed in data" [86]. 1791 1792 1793 1794 1795 1796 1797 Strike a balance, share decisions: Catherine Nicole 1798 Coleman suggests the information sciences have had to 1799 approach with a perspective of managing rather than 1800 eliminating classification and bias, by grappling with 1801 the dynamic balance between curating information and 1802 sharing it. It relies on such decisions being made over 1803 time and distributed among diverse groups [87]. 1804 1805 1806 1807 1808 1809 1810 Ask essential questions: Even before deciding whether an 1811 algorithm is the answer to a problem, technologist Kamal 1812 Sinclar recommends asking, "Can the available data lead 1813 to a good outcome?" and "Will the people affected by 1814 these decisions have any influence over the system?" 1815 [88] These are the kinds of questions this guide tries 1816 to parse out in detail for each stage of the dataset 1817 lifecycle. 1818 1819 1820 1821 1822 1823 8 1824 CONCLUSION 1825 1826 1827 1828 1829 1830 1831 1832 1833 1834 1835 1836 1837 1838 1839 1840 1841 1842 1843 1844 1845 1846 1847 1848 THIS GUIDE AIMS TO HELP YOU WORK CRITICALLY WITH MACHINE LEARNING DATASETS: 1849 1850 1851 - to see existing datasets from different perspectives; 1852 - to appreciate the complexities in their origins, 1853 classifications, and transformations; 1854 - to read the messiness across the lifecycles of datasets; 1855 - to reach out to people impacted by your dataset work; - and, perhaps most importantly, to understand the benefits of 1856 advancing your projects with thoughtful, accountable data stewardship. 1857 We hope that you dip into it, revisit parts, follow-up on references, and share pieces that you think are helpful to your teams and communities. Given the recent proliferation of text-to-image machine learning and synthetic data, and no doubt new tools and applications all the time, we hope you also use 1858 the critical perspectives and guidelines you develop here with new technologies as they emerge. We intended this guide not as a definitive source -- many of these topics are too big to be covered completely -- but as a starting point for your explorations and an illustration of how productive and fruitful it can be to approach dataset work critically. 1859 We also hope the guide is a prompt for more interdisciplinary and intersectional conversations about critical dataset work -- and the promise of conscientious approaches to machine learning. 1860 With continued efforts, our hope is for this field guide to be a living document with expansions, updates, and additional resources as the dynamic world of machine learning datasets continues to evolve. 1861 1862 1863 - 1864 CREDITS, ACKNOWLEDGMENTS AND DOWNLOAD 1865 1866 1867 1868 1869 1870 1871 1872 1873 1874 1875 1876 1877 1878 1879 1880 1881 1882 1883 1884 1885 1886 1887 1888 CREDITS 1889 Author: Sarah Ciston Editors: Mike Ananny and Kate Crawford Design and illustrations: Vladan Joler and Olivia Solis 1890 Published by: Knowing Machines project (https:// knowingmachines.org) Full citation: S. Ciston, "A CRITICAL FIELD GUIDE FOR WORKING WITH MACHINE LEARNING DATASETS," K. Crawford and M. Ananny, Eds., Knowing Machines project, Feb. 2023. We wish to thank all the members of the Knowing Machines research project, including Christo Buschek, Franny Corry, Melodi Dincer, Vladan Joler, Ed Kang, Sasha Luccioni, Will Orr, Jason Schultz, Hamsini Sridharan, and Jer Thorp for their fruitful conversations and generous feedback on drafts of this work, and Hannah Franklin and Michael Weinberg for their administrative support. Warm thanks go to Lee Kezar for their rigorous technical perspective and contributions, and also much appreciation to Will Orr for proofreading, Vladan Joler and 1891 Olivia Solis for design, and Michael Weinberg for project management. We also want to acknowledge the members of the USC Libraries' "Visual Datasets for Inclusive Research" project, including Hujefa Ali, Bill Dotson, Curtis Fletcher, Mike Jones, Caroline Muglia, and Manasa Rajesh for providing inspiration for this work, valuable perspectives on library collections as data, and a warm and enriching research environment. We want to acknowledge the support of the Alfred. P Sloan Foundation, as part of their funding of the Knowing Machines project. Finally, the author wishes to thank editors Kate Crawford and Mike Ananny for their consistently kind and thoughtful support throughout. 1892 1893 DESIGN NOTE 1894 In the tradition of the early net.art experimentation, this Guide was entirely created within a spreadsheet. This experimental design concept is exploring possibilities and constraints of the spreadsheet as a medium that plays an 1895 important role in the creation of the machine learning datasets. The illustrations, inspired by early modernism and optical art, play with the idea of a "bureaucratic modernism"-style of art, fitting for an age where everyone is expected to take on the roles of both manager and bureaucrat. 1896 1897 ABOUT KNOWING MACHINES 1898 This critical field guide is published by Knowing Machines. Knowing Machines is a research project tracing the histories, practices, and politics of how machine learning systems are trained to interpret the world. Our group develops methodologies and tools for understanding, analyzing, and investigating training datasets, and studying 1899 their role in the construction of "ground truth" for machine learning. We research how datasets index the world, make predictions, and structure knowledge cultures. We are an international team, and we aim to support the emerging field of critical data studies by contributing original research, reading lists, research tools, and supporting communities of inquiry that address the foundational epistemologies of machine learning. Knowing Machines is sponsored by the Alfred P. Sloan Foundation. 1896 1897 DOWNLOAD 1898 PDF version: https://knowingmachines.org/docs/ 1899 critical_field_guide.pdf Original spreadsheet version: https://bit.ly/criticalfieldguide 1900 1901 1902 REFERENCES 1903 1904 [1] Y. A. Loukissas, All Data Are Local: Thinking Critically 1905 in a Data-Driven Society. Cambridge: MIT Press, The MIT Press, 2019. 1906 [2] C. D'Ignazio and L. F. Klein, Data Feminism. Cambridge, MA, USA: MIT Press, 2020. L. Poirier, "Reading datasets: Strategies for 1907 [3] interpreting the politics of data signification," Big Data Soc., vol. 8, no. 2, p. 20539517211029320, Jul. 2021, doi: 10.1177/20539517211029322. [4] S. Browne, Dark matters: on the surveillance of 1908 blackness. Durham, [North Carolina] ; Duke University Press, 2015. N. Couldry and U. A. Mejias, "Making data colonialism liveable: how might data's social order be regulated?," 1909 [5] Internet Policy Rev., vol. 8, no. 2, Jun. 2019, Accessed: Mar. 21, 2021. https://policyreview.info/ articles/analysis/ making-data-colonialism-liveable-how-might-datas-social-order-be-regulated J. Angwin, J. Larson, S. Mattu, and L. Kirchner, 1910 [6] "Machine Bias," ProPublica, May 2016, Accessed: Apr. 27, 2019. https://www.propublica.org/article/ machine-bias-risk-assessments-in-criminal-sentencing D. Wagner, "17. Classification -- Computational and Inferential Thinking," in Computational and Inferential 1911 [7] Thinking: The Foundations of Data Science, 2nd ed., A. Adhikari, J. DeNero, and D. Wagner, Eds. Accessed: Nov. 25, 2022. https://computerscience.chemeketa.edu/ datasci-text/chapters/17/Classification.html [8] S. Leonelli, Data-Centric Biology: A Philosophical 1912 Study. University of Chicago Press, 2016. doi: 10.7208/ chicago/9780226416502.001.0001. [9] T. Gillespie, "The Relevance of Algorithms," T. 1913 Gillespie, P. J. Boczkowski, and K. A. Foot, Eds. Cambridge, MA: MIT Press, 2014. 1914 [10] L. Gitelman, Ed., "Raw data" is an oxymoron. Cambridge, Massachusetts ; London, England: The MIT Press, 2013. [11] C. Koopman, How We Became Our Data: A Genealogy of the 1915 Informational Person. University of Chicago Press, 2019. doi: 10.7208/9780226626611. [12] C. L. Borgman, Big Data, Little Data, No Data: 1916 Scholarship in the Networked World. 2015. doi: 10.7551/ mitpress/9963.001.0001. 1917 [13] R. Kitchin, Data Lives. Policy Press, 2021. [14] J. Gleick, The information: a history, a theory, a 1918 flood, 1st Vintage Books ed., 2012. New York: Vintage Books, 2011. N. B. Thylstrup, D. Agostinho, A. Ring, C. D'Ignazio, [15] and K. Veel, Eds., Uncertain Archives: Critical Keywords 1919 for Big Data. 2021. Accessed: Mar. 29, 2021. [Online]. Available: https://doi.org/10.7551/mitpress/ 12236.001.0001 [16] T. Striphas, "Algorithmic culture," Eur. J. Cult. Stud., 1920 vol. 18, no. 4-5, pp. 395-412, Aug. 2015, doi: 10.1177/ 1367549415577392. [17] L. Datson, Rules. Princeton University Press, 2022. 1921 https://press.princeton.edu/books/hardcover/ 9780691156989/rules 1922 [18] F. Chollet, Deep Learning with Python, Second Edition. New York: Manning Publications Co. LLC, 2021. W. Van Rossem and A. Pelizza, "The ontology explorer: A [19] method to make visible data infrastructures for 1923 population management," Big Data Soc., vol. 9, no. 1, p. 20539517221104090, Jan. 2022, doi: 10.1177/ 20539517221104087. "Lawsuits allege Microsoft, Amazon and Google violated 1924 [20] Illinois facial recognition privacy law," TechCrunch. https://social.techcrunch.com/2020/07/15/ facial-recognition-lawsuit-vance-janecyk-bipa/ . "Facial recognition's 'dirty little secret': Social 1925 [21] media photos used without consent," NBC News. https:// www.nbcnews.com/tech/internet/ facial-recognition-s-dirty-little-secret-millions-online-photos-scraped-n981921 F. Corry, E. B. Kang, H. Sridharan, S. Luccioni, M. 1926 [22] Ananny, and K. Crawford, "Critical Dataset Studies Reading List," Knowing Machines. https:// knowingmachines.org/reading-list [23] Y. Gil, "Yolanda Gil: Teaching Data Science to 1927 Non-Programmers," Jan. 03, 2020. https://www.isi.edu/ ~gil/teaching/TeachingDataScienceToNonProgrammers.html . [24] T. Gebru et al., "Datasheets for Datasets," 1928 ArXiv180309010 Cs, Mar. 2020, Accessed: Apr. 01, 2021. http://arxiv.org/abs/1803.09010 [25] K. Crawford, Atlas of AI: power, politics, and the 1929 planetary costs of artificial intelligence. New Haven: Yale University Press, 2021. 1930 [26] F. Cady, The data science handbook. Hoboken, NJ: John Wiley & Sons, Inc., 2017. [27] B. Friedman and D. G. Hendry, Value Sensitive Design: 1931 Shaping Technology with Moral Imagination. 2019. doi: 10.7551/mitpress/7585.001.0001. "Using artificial intelligence, geo-journalism and data journalism, journalists dodge some of the dangers of [28] covering the Amazon," LatAm Journalism Review by the 1932 Knight Center, Jul. 19, 2022. https:// latamjournalismreview.org/articles/ using-artificial-intelligence-geo-journalism-and-data-journalism-journalists-dodge-some-of-the-dangers-of-covering-the-amazon / . 1933 [29] "PHI Greek Inscriptions." https:// inscriptions.packhum.org/ . [30] "Ithaca | Restoring and attributing ancient texts using 1934 deep neural networks," Ithaca. https:// ithaca.deepmind.com . P. Oliveira, "On the Endless Infrastructural Reach of a 1935 [31] Phoneme," Transmedialeart Digit. Cult., no. 3, https:// archive.transmediale.de/content/ on-the-endless-infrastructural-reach-of-a-phoneme [32] Canavan, Alexandra and Zipperlen, George, "CALLFRIEND 1936 Egyptian Arabic." Linguistic Data Consortium, p. 1401448 KB, 1996. doi: 10.35111/NNM5-KP69. Canavan, Alexandra, Zipperlen, George, and Graff, David, 1937 [33] "CALLHOME Egyptian Arabic Speech." Linguistic Data Consortium, p. 1807744 KB, 1997. doi: 10.35111/ D8YB-9M13. A. Biselli, "Eine Software des BAMF bringt Menschen in [34] Gefahr," Vice, Aug. 20, 2018. https://www.vice.com/de/ 1938 article/a3q8wj/ fluechtlinge-bamf-sprachanalyse-software-entscheidet-asyl . 1939 [35] P. Oliveira, "Personal conversation," Aug. 16, 2022. 1940 [36] P. Oliveira, "On The Apparently Meaningless Texture of Noise." http://meaninglesstexture.schloss-post.com/ . M. Murgia, "Researchers train AI on 'synthetic data' to 1941 [37] uncover Syrian war crimes," FT.com, Dec. 2021, Ahttps:// www.proquest.com/docview/2617702310/citation/ 42933059DCF5492DPQ/1 "AI Emerges as Crucial Tool for Groups Seeking Justice [38] for Syria War Crimes," Dow Jones Institutional News, Dow 1942 Jones & Company Inc, New York, United States, Feb. 13, 2021. http://www.proquest.com/docview/2489065006/ citation/C2B3DB568C8D4995PQ/1 1943 [39] "Data Science 4 All - a user friendly data science learning site." https://datascience4all.org/ . 1944 [40] W. McKinney, Python for Data Analysis, 3E, Open 3rd Edition. O'Reilly. https://wesmckinney.com/book/ [41] Cleaning Data for Effective Data Science. Accessed: Jul. 1945 31, 2022. https://learning.oreilly.com/library/view/ cleaning-data-for/9781801071291/ [42] K. Kuksenok, "Consider Data Cleaning v1.1," presented at 1946 the Resistance AI Workshop at NeurIPS2020, Nov. 2020. https://ksen0.github.io/code-data-work/ [43] T. McPherson, Feminist in a Software Lab: Difference + 1947 Design. Cambridge, Massachusetts ; London, England: Harvard University Press, 2018. [44] M. Onuoha, "The Library of Missing Datasets -- MIMI 1948 ONUOHA," MIMI ONUOHA. https://mimionuoha.com/ the-library-of-missing-datasets K. Crenshaw, "Demarginalizing the Intersection of Race 1949 [45] and Sex: A Black Feminist Critique of Antidiscrimination Doctrine, Feminist Theory and Antiracist Politics," Univ. Chic. Leg. Forum, vol. 1989, pp. 139-168, 1989. [46] M. Onuoha, "The Point of Collection," Medium, Oct. 31, 1950 2016. https://points.datasociety.net/ the-point-of-collection-8ee44ad7c2fa . [47] "Responsible Data Handbook | Getting Data." https:// 1951 the-engine-room.github.io/responsible-data-handbook/ chapters/chapter-02a-getting-data.html . 1952 [48] The Engine Room, "Responsible Data Handbook." https:// the-engine-room.github.io/responsible-data-handbook/ . E. Denton, M. Diaz, I. Kivlichan, V. Prabhakaran, and R. [49] Rosen, "Whose Ground Truth? Accounting for Individual 1953 and Collective Identities Underlying Dataset Annotation," arXiv, arXiv:2112.04554, Dec. 2021. doi: 10.48550/arXiv.2112.04554. A. S. Luccioni, F. Corry, H. Sridharan, M. Ananny, J. Schultz, and K. Crawford, "A Framework for Deprecating 1954 [50] Datasets: Standardizing Documentation, Identification, and Communication," in 2022 ACM Conference on Fairness, Accountability, and Transparency, New York, NY, USA, Jun. 2022, pp. 199-212. doi: 10.1145/3531146.3533086. E. Denton, M. Diaz, I. Kivlichan, V. Prabhakaran, and R. [49] Rosen, "Whose Ground Truth? Accounting for Individual 1955 and Collective Identities Underlying Dataset Annotation," arXiv, arXiv:2112.04554, Dec. 2021. doi: 10.48550/arXiv.2112.04554. A. S. Luccioni, F. Corry, H. Sridharan, M. Ananny, J. Schultz, and K. Crawford, "A Framework for Deprecating 1956 [50] Datasets: Standardizing Documentation, Identification, and Communication," in 2022 ACM Conference on Fairness, Accountability, and Transparency, New York, NY, USA, Jun. 2022, pp. 199-212. doi: 10.1145/3531146.3533086. 1957 [51] Design Justice Network, "Design Justice for Action," Des. Justice Zines, no. #3. [52] M. Zook et al., "Ten simple rules for responsible big 1958 data research," PLOS Comput. Biol., vol. 13, no. 3, p. e1005399, Mar. 2017, doi: 10.1371/journal.pcbi.1005399. [53] "CARE Principles of Indigenous Data Governance," Global 1959 Indigenous Data Alliance. https://www.gida-global.org/ care . 1960 [54] Open Data Institute, "The Data Ethics Canvas." https:// theodi.org/article/the-data-ethics-canvas-2021/ . 1961 [55] "Local Contexts - Grounding Indigenous Rights." https:// localcontexts.org/ . 1962 [56] "FAIR Principles," GO FAIR. https://www.go-fair.org/ fair-principles/ . "CARE Principles of Indigenous Data Governance," Global 1963 [57] Indigenous Data Alliance. https://www.gida-global.org/ care . S. Barocas, K. Crawford, A. Shapiro, and H. Wallach, 1964 [58] "The Problem With Bias: Allocative Versus Representational Harms in Machine Learning.," in Proceedings of SIGCIS, Philadelphia, PA, 2017. [59] R. Benjamin, Race After Technology: Abolitionist Tools 1965 for the New Jim Code, 1 edition. Medford, MA: Polity, 2019. A. Paullada, I. D. Raji, E. M. Bender, E. Denton, and A. [60] Hanna, "Data and its (dis)contents: A survey of dataset 1966 development and use in machine learning research," Patterns, vol. 2, no. 11, p. 100336, Nov. 2021, doi: 10.1016/j.patter.2021.100336. H. Suresh and J. V. Guttag, "A Framework for [61] Understanding Sources of Harm throughout the Machine 1967 Learning Life Cycle," in Equity and Access in Algorithms, Mechanisms, and Optimization, Oct. 2021, pp. 1-9. doi: 10.1145/3465416.3483305. [62] M. Miceli and J. Posada, "The Data-Production 1968 Dispositif." arXiv, May 24, 2022. doi: 10.48550/ arXiv.2205.11963. S. Hong, "Prediction as Extraction of Discretion," in 1969 [63] 2022 ACM Conference on Fairness, Accountability, and Transparency, New York, NY, USA, Jun. 2022, pp. 925-934. doi: 10.1145/3531146.3533155. [64] C. Apprich, W. H. K. Chun, F. Cramer, and H. Steyerl, 1970 Pattern Discrimination. Minneapolis: University of Minnesota Press, 2018. O. Keyes and J. Austin, "Feeling fixes: Mess and emotion 1971 [65] in algorithmic audits," Big Data Soc., vol. 9, no. 2, p. 20539517221113772, Jul. 2022, doi: 10.1177/ 20539517221113772. S. L. Blodgett, S. Barocas, H. Daume III, and H. 1972 [66] Wallach, "Language (Technology) is Power: A Critical Survey of 'Bias' in NLP." arXiv, May 29, 2020. doi: 10.48550/arXiv.2005.14050. [67] G. C. Bowker and S. L. Star, Sorting things out: 1973 classification and its consequences. Cambridge, Mass: MIT Press, 1999. [68] V. U. Prabhu and A. Birhane, "Large image datasets: A 1974 pyrrhic win for computer vision?" arXiv, Jul. 23, 2020. doi: 10.48550/arXiv.2006.16923. [69] A. Blair, P. Duguid, A.-S. Goeing, and A. Grafton, 1975 Information: A Historical Companion. Princeton University Press, 2021. doi: 10.1515/9780691209746. [70] H. A. Olson, "The Power to Name: Representation in 1976 Library Catalogs," SignsJournal Women Cult. Soc., vol. 26, no. 3, 2001, doi: 10.1086/495624. T. Gillespie and N. Seaver, "Critical Algorithm Studies: 1977 [71] a Reading List," Social Media Collective, Nov. 05, 2015. https://socialmediacollective.org/reading-lists/ critical-algorithm-studies/ . B. Hutchinson and M. Mitchell, "50 Years of Test (Un) [72] fairness: Lessons for Machine Learning," in Proceedings 1978 of the Conference on Fairness, Accountability, and Transparency, Atlanta GA USA, Jan. 2019, pp. 49-58. doi: 10.1145/3287560.3287600. J. Phang, A. Chen, W. Huang, and S. R. Bowman, 1979 [73] "Adversarially Constructed Evaluation Sets Are More Challenging, but May Not Be Fair." arXiv, Nov. 15, 2021.http://arxiv.org/abs/2111.08181 N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, and A. 1980 [74] Galstyan, "A Survey on Bias and Fairness in Machine Learning." arXiv, Jan. 25, 2022. http://arxiv.org/abs/ 1908.09635 [75] D. Hupkes et al., "State-of-the-art generalisation 1981 research in NLP: a taxonomy and review." arXiv, Oct. 10, 2022. http://arxiv.org/abs/2210.03050 X. Bai et al., "Explainable deep learning for efficient 1982 [76] and robust pattern recognition: A survey of recent developments," Pattern Recognit., vol. 120, p. 108102, Dec. 2021, doi: 10.1016/j.patcog.2021.108102. [77] K. Crenshaw, What Does Intersectionality Mean? : 1A. 1983 2021. https://www.npr.org/2021/03/29/982357959/ what-does-intersectionality-mean K. Crenshaw, "Demarginalizing the Intersection of Race [78] and Sex: A Black Feminist Critique of Antidiscrimination 1984 Doctrine, Feminist Theory and Antiracist Politics," Univ. Chic. Leg. Forum, vol. 1989, no. 1, Dec. 2015, https://chicagounbound.uchicago.edu/uclf/vol1989/iss1/8 B. Cooper, "Intersectionality," in The Oxford Handbook 1985 [79] of Feminist Theory, vol. 1, L. Disch and M. Hawkesworth, Eds. Oxford University Press, 2016. doi: 10.1093/ oxfordhb/9780199328581.013.20. B. Gipson, F. Corry, and S. U. Noble, 1986 [80] "Intersectionality," in Uncertain Archives: Critical Keywords for Big Data, 2021. https://doi.org/10.7551/ mitpress/12236.003.0027 [81] S. U. Noble and B. M. Tynes, Eds., The intersectional 1987 Internet : race, sex, class and culture online. New York: Peter Lang Publishing, Inc, 2016. 1988 [82] S. Ciston, "Intersectional AI Toolkit," Intersectional AI Toolkit. https://intersectionalai.com/ A. Wang, V. V. Ramaswamy, and O. Russakovsky, "Towards Intersectionality in Machine Learning: Including More 1989 [83] Identities, Handling Underrepresentation, and Performing Evaluation," in 2022 ACM Conference on Fairness, Accountability, and Transparency, New York, NY, USA, Jun. 2022, pp. 336-349. doi: 10.1145/3531146.3533101. P. Kalluri, "Don't ask if artificial intelligence is 1990 [84] good or fair, ask how it shifts power," Nature, vol. 583, no. 7815, Art. no. 7815, Jul. 2020, doi: 10.1038/ d41586-020-02003-2. A. L. Hoffmann, "Data Violence and How Bad Engineering [85] Choices Can Damage Society," Medium, Apr. 30, 2018. 1991 https://medium.com/s/story/ data-violence-and-how-bad-engineering-choices-can-damage-society-39e44150e1d4 . M. Miceli, J. Posada, and T. Yang, "Studying Up Machine 1992 [86] Learning Data: Why Talk About Bias When We Mean Power?," Proc. ACM Hum.-Comput. Interact., vol. 6, no. GROUP, p. 34:1-34:14, Jan. 2022, doi: 10.1145/3492853. [87] C. N. Coleman, "Managing Bias When Library Collections 1993 Become Data," Int. J. Librariansh., vol. 5, no. 1, Art. no. 1, Jul. 2020, doi: 10.23974/ijol.2020.vol5.1.162. [88] K. Sinclair and J. Clark, Making a New Reality. 2020. 1994 https://makinganewreality.org/ making-a-new-reality-a-toolkit-for-inclusive-media-futures-a3bdc0e68f20 1995 [89] "Hugging Face." https://huggingface.co/datasets . 1996 [90] "Kaggle." https://www.kaggle.com/datasets . 1997 [91] "Papers With Code." https://paperswithcode.com/datasets . 1998 [92] "re3data.org." https://www.re3data.org/ . 1999 [93] "Zenodo." https://zenodo.org/ . M. Khan and A. Hanna, "The Subjects and Stages of AI 2000 [94] Dataset Development: A Framework for Dataset Accountability." Rochester, NY, Sep. 13, 2022. Accessed: Sep. 19, 2022. https://papers.ssrn.com/abstract=4217148 2001 2002 2003 ENDNOTES 2004 2005 2006 {1}| [email protected], University of Southern California 2007 {2}| [email protected], University of Southern California 2008 {3}| [email protected], University of Southern California, MSR-NYC This field guide encourages a combination of criticality and care toward datasets -- plus machine learning materials and processes more widely, and the people and environments they impact -- so that the meaning of 'data stewardship' resonates with its connotations of environmental stewardship and conservation, as well as its definitions of technical responsibility. Media scholar Yanni Alexander Loukissas lays out this approach in his book All Data Are Local. He says, "I take a 2009 {4}| critical stance, but also explore approaches to working with data that are less distant and cerebral than critical reflection implies. In order to do so, my approach integrates lessons from the feminist ethics of care. [...] Care is critical in that it calls attention to neglected things. But it is more than critical reflection; it is a doing practice" [1]. As Catherine D'Ignazio and Lauren Klein suggest in Data Feminism, this approach emphasizes embodiment and material contexts, with the potential to rethink hierarchies and challenge power [2]. There is a growing body of scholarship emphasizing the need to approach datasets critically. For recent research on this topic, see cultural media theorist Nanna Bonde Thylstrup's introduction to "critical dataset studies" [3] and the dataset accountability 2010 {5}| framework and matrix of dataset development harms from Yale Law resident fellow Mehtab Khan's and Distributed AI Research Institute research director Alex Hanna [94]. Find more resources on the "Critical Dataset Studies Reading List" compiled by the Knowing Machines research project. That the language of big data is reminiscent of global colonialist exploitation ("scrape," "extract," "capture," "the new oil") should be telling. Work by 2011 {6}| critical surveillance scholar Simone Browne and by Nick Couldry and Ulises A. Mejias, among others, has already traced the colonialist legacies big data builds upon [4] , [5]. For more on the sociotechnical phenomenon, see CLASSIFICATION THINKING. For more on machine learning 2012 {7}| definitions of classification tasks, as well as a technical introduction to data science concepts, see Computational and Inferential Thinking, edited by statistics professor Ani Adhikari, et al. [7]. 'Data' and 'information' are complex terms. There is a healthy scholarly conversation about the many meanings of each word, beyond the scope of this guide, including Tarleton Gillespie's "The Relevance of Algorithms" [9]; Lisa Gitelman's edited volume "Raw Data" Is an Oxymoron [10], Colin Koopman's How We Became Our Data [11]; 2013 {8}| Christine L. Borgman's Big Data, Little Data, No Data [12]; and Rob Kitchin's Data Lives [13]. For more on information, see Ann Blair, et al.,'s Information: A Historical Companion and James Gleik's The Information: A History, A Theory, A Flood [14]; and for a detailed description of additional terms, see the anthology Uncertain Archives: Critical Keywords for Big Data [15]. TOOLS OF THE TRADE: If you're working in the PYTHON programming language, you might use two popular tools called NUMPY and PANDAS to make these adjustments. They are both LIBRARIES or MODULES, which are add-on packs of software available to import. Often created with a specific field or task in mind, libraries are written on 2014 {9}| top of Python so that you don't have to write everything you want to do from scratch. Numpy helps with handling large groups of numbers, and Pandas has built-in support for lots of data manipulation tasks. You may also encounter the MATPLOTLIB library for making visualizations and SCIKIT-LEARN, KERAS, or other machine learning libraries down the line. Popular dataset repositories include Hugging Face, 2015 {10}| Kaggle, Papers with Code, the Registry of Research Data Repositories, and Zenodo [89]-[93]. MORE PITFALLS: MIT researchers Harini Suresh and John Guttag break data representation down further into seven types of harm, which they refer to as "bias," encountered across machine learning processes. These include historical: e.g. word embeddings that reflect and reinforce stereotypes; representational: e.g. underrepresenting or misrepresenting the target group, through limited sampling or a mismatch between target and use populations; measurement: e.g. variations in 2016 {11}| accuracy or method across groups; learning: e.g. pruning the data to enhance performance ends up amplifying disparities on underrepresented characteristics; evaluation: e.g. comparison against standardized benchmarks fails to detect issues when the benchmarks themselves are also biased; aggregation: e.g. applying an overgeneralized assumption to an entire set when subsets should be represented differently; and deployment: e.g. misalignment of how a dataset or model was designed and how it is used in practice [61]. EXPLAINING & DESCRIBING, OR EXTRACTING & PRESCRIBING: Prediction, says communications scholar Sun-Ha Hong, "sees what it knows to see, and it measures what it can typically imagine measuring. These tendencies are shaped through longstanding economic and political asymmetries, whose influence is regularly written off as uncertain and uncontrollable 'externalities'. [...T]hey obfuscate how patterns of extraction shape the research questions 2017 {12}| and the choice of what to measure (and what to dismiss without measuring)" [63]. In their work on data colonialism, Nick Couldry and Ulises A. Mejias show that digital datasets join a much longer history of extraction [5]. As digital media researcher Wendy H.K. Chun argues, the materials and methods of machine learning -- including datasets -- work by forecasting the future from past data and prescribing what they purport to describe [64]. For a thoughtful discussion of how datasets become recontextualized -- and in particular the importance of using critical and historical analyses to question the 2018 {13}| continued circulation of datasets in communities not accustomed to their original contexts -- see research on dataset audits by critical technology researcher Os Keyes and librarian Jeanie Austin [65]. CLASSIFICATION THINKING: For multifaceted perspectives on the logics and politics of classification, please see, among many others: In Sorting Things Out [67], informatics professor Geoffrey C. Bowker and sociologist Susan C. Starr suggest that even if categories often feel invisible, "The material force of categories appears always and instantly." Creating a category draws a boundary and fixes the concept of what is inside and outside, as seen from the perspective of whomever has the power to create it. "Categories simplify and freeze nuanced and complex narratives, obscuring political and moral reasoning behind a category," argue computer scientists Vinay Uday Prabhu and Abeba Birhane [68]. Library scholar Hope A. Olson points out that classification problems are not new to the machine learning field; rather, the problematic goal to find "an overriding unity in language" has been codified through library catalog practices since at least the nineteenth 2019 {14}| century. It echoes Enlightenment-era impulses to "know" the world comprehensively [69], [70]. Furthermore, algorithmic attempts to understand people through classification are drawing on much longer practices of human exploitation that have created and justified categories of difference. In Dark Matters, critical surveillance scholar Simone Browne traces data practices like surveillance and biometrics back to the documents of the transatlantic slave trade, arguing that, "human categorization and division is part of a larger imperial project of colonial expansion that aimed to fix, frame, and naturalize discursively constructed difference" [4]. These are just a few (non-exhaustive) touch points for the discussion around classification thinking and machine learning. For more, you can look to the "Critical Dataset Studies Reading List," compiled by the Knowing Machines research project [22], or the "Critical Algorithm Studies: a Reading List" curated by Tarleton Gillespie and Nick Seaver [71]. A meta analysis by Dieuwke Hupke et al. found that a majority (66%) of efforts focused on "practical" improvements, while only 2.6% focus on fairness. This included generalizability, considering in what kinds of situations a model can be applied: "One question that is often addressed with a primarily practical motivation is 2020 {15}| how well models generalise to different domains or differently collected data." Meanwhile, fairness research "asks questions about how well models generalise to diverse demographics, typically considering minority or marginalised groups [...] or investigates to what extent models perpetuate (undesirable) biases learned from their training data" [75]. 2021