https://www.theregister.com/2022/04/18/fake_ai_data/ [user] [user] Sign in The Register(r) -- Biting the hand that feeds IT [magn] [burg] [burg] Topics Security Off-Prem All Off-PremEdge + IoTChannelPaaS + IaaSSaaS (X) On-Prem All On-PremSystemsStorageNetworksHPCPersonal Tech (X) Software All SoftwareAI + MLApplicationsDatabasesDevOpsOSesVirtualization (X) Offbeat All OffbeatDebatesColumnistsScienceGeek's GuideBOFHLegalBootnotesSite NewsAbout Us (X) Vendor Voice All Vendor VoiceAdobeAmazon Web Services (AWS) MigrationCofense EchoworxGoogle CloudGoogle Cloud's ApigeeGoogle WorkspaceNutanix Rapid7SophosVeeam (X) Resources * Whitepapers * Webinars * Newsletters Situation Publishing * The Next Platform * Devclass * Blocks and Files Get our Weekly newsletter [aiml] AI + ML Fake it until you make it: Can synthetic data help train your AI model? Yes and no. It's complicated. Katyanna Quach Mon 18 Apr 2022 // 11:33 UTC 2 comment bubble on white --------------------------------------------------------------------- 2 comment bubble on white # reddit Twitter Facebook linkedin WhatsApp email [https://www.theregis] Copy The saying "data is the new oil," was reportedly coined by British mathematician and marketing whiz Clive Humby in 2006. Humby's remark rings true more now than ever with the rise of deep learning. Data is the fuel powering modern AI models; without enough of it the performance of these systems will sputter and fail. And like oil, the resource is scarce and controlled by big businesses. What do you do if you're a small computer vision company? You can turn to fake data to train your models, and if you're lucky it might just work. The market for synthetic data generation grew to over $110 million in 2021 and is expected to increase to $1.15 billion by the end of 2027, according to a report published by research firm Cognilytica. [aiml] Numerous startups have built tools to spin up synthetic images to help companies train their machine learning algorithms. [aiml] [aiml] There are many benefits to using computer-generated data, Gil Elbaz, co-founder and CTO of Datagen, explained to The Register. The startup, founded in 2018 and based in Israel, has built a software platform that allows customers to easily create mock images at the click of a button. Synthetic data provides a way to scale up datasets and automatically annotate each picture with the necessary metadata without much human labor. [aiml] Issues of privacy and bias can be avoided too. "Privacy for human faces is very, very hard, and it's not ideal to even hold [that kind of data] in your servers," Elbaz says. "With our data, there's no [personally identifiable information]. This is not a real person. This is completely synthetic, so there's no privacy issues. And bias-wise we can generate whatever distribution of ethnicities, ages, genders you want in your data, so we are not biased in any way," he says as he shows us a three-dimensional fake face. * Machine learning models leak personal info if training data is compromised * AI-created faces now look so real, humans can't spot the difference * Cerebras' wafer-size AI chips play nice with PyTorch, TensorFlow * US military wants $29.8m for IT to boost AI intel analysis Datagen works with companies to train computer vision models for different tasks. Simulated data is used by the automotive industry to develop AI software that automatically detects driver behavior, such as when they're distracted or falling asleep at the wheel. Fake data has also been used by surveillance camera companies to flag whenever packages have been delivered outside people's homes. AI applications in augmented and virtual reality also benefit from ingesting copious amounts of synthetic data. Rendering fake data is a complicated process. Datagen uses multiple methods to create computer-made images, from physics-based ray tracing algorithms to generative adversarial networks (GANs). Making the data is the easy part. Getting a model trained on false images to work in the real world is the challenge. Ideally, companies should have some real data to hand and can't just rely on fake data. [aiml] "What we see working really well is to train the network on a large amount of synthetic data, and then fine tune it on the small amount of real data. This last step is optional. "It's really not a must, but it does improve the performance to do a small fine tune on the real world. What this means in practice is that you need much less real world data. So you don't need as much, you can use like 1/20th or 1/50th of the amount of real data and use mostly synthetic data for your training," Elbaz says. Models trained on fake images have to be robust enough to work in real-life settings. Synthetic data has been successful in training self-driving cars to recognise things like cars, road signs, and pedestrians in its environment and simulate driving the same roads in different weather conditions. It has proven useful in robotics too in limited scenarios, like getting mechanical grippers to rotate or pick up objects. Simulation to reality Developers relying on synthetic data have to test and tweak their models rigorously to make sure they'll work. "If you test your models in a good way, the idea is that your test should validate that the performance will be of high quality or have a quality that you expect. If your testing is not as good or if you don't have enough test data, then you can find a gap in the performance," says Elbaz. "We can do testing to see where the neural network is weak, pretty much by trying to ask it, for example, what do you think about this guy? And if I make him darker, or if I make him further away, or if I change him to look more angry? What do you think about that? And I can ask the network all of these different things and see where it's weaker, and really map out the weaknesses of the network itself," Elbaz says. * AI pioneer suggests trickle-down approach to machine learning * 'Virtually no difference' between AI and humans in diagnosing prediabetes * How Nvidia is overcoming slowdown issues in GPU clusters * Microsoft to upgrade language translator with new class of AI model But in some cases the real world is too difficult to model, and synthesizing data samples won't be worthwhile. "There is a very high effort that's required in order to build [for niche things]. Say, if you're trying to understand where a dog's nose is in an image, we don't do synthetic data for dog noses. Trying to pick something like that out on your own is extremely hard." These gaps open up opportunities for startups that use synthetic data in a different way. Synthetaic, based in Wisconsin and founded in 2019 by Corey Jaskolski, doesn't sell computer-generated images to customers. Instead, it uses generative models like GANs or transformers to help image detection algorithms automatically label objects. "We're still building AI that is capable of generating synthetic data. However, the novel piece that we're doing is we're not using it to generate synthetic data to then use to train an AI. We're using this generative capability to create, effectively, a way to look at real world data that allows us to do things like this auto-labeling," Jaskolski tells The Register. "What's going on behind the scenes here is using a transformer technology that is usually used to generate imagery, but because it's so powerful and good at generating imagery, it's actually also so powerful and good at describing real world imagery in a way that lets you click on a single image and [detect others like it]." Synthetaic showed El Reg a demo, where its Rapid Automatic Image Categorization (RAIC) technology could zero in on specific frames in a video feed. Jaskolski fed the system a photograph of a cheetah, and RAIC was able to find instances where a cheetah popped up in the video. Real is always better Real data is still more important for Synthetaic, despite the company's somewhat confusing name. "There are lots of examples in defense and in other industry applications where just adding 3D data or synthetic data doesn't fix the problem. I think because every situation is different, and AI always has trouble transferring from domains. It might not transfer well to the real world." Generating synthetic data is a great way to create a larger and more diverse dataset, but it's only effective for training machine learning algorithms that perform jobs that aren't too simple and aren't too complex either. Easy computer vision tasks doesn't always require fake data, and AI. Difficult tasks require a high level of detail in simulated images and expert knowledge is needed to assess its quality. "I think that medical data is a really good example of a use case that we don't want to work on," says Elbaz. "In order to model medical diseases, you need real doctors to help you. "There's a lot of specialized knowledge that you would need in order to create this medical synthetic data. Even though medical data is extremely valuable. It's something that I think requires a separate company. It's just too hard. Anything that requires very, very, specialized knowledge is hard," he concluded. (r) Get our Tech Resources # Share reddit Twitter Facebook linkedin WhatsApp email [https://www.theregis] Copy 2 Comments Similar topics * AI * Data Broader topics * Self-driving Car Narrower topics * Google AI * GPT-3 * Machine Learning * MCubed * Tensor Processing Unit Corrections Send us news --------------------------------------------------------------------- [aiml] Other stories you might like * Uncle Sam probes Activision for any insider trading Troubled games maker reveals investigation amid Microsoft takeover Katyanna Quach Mon 18 Apr 2022 // 21:34 UTC comment bubble on black Activision Blizzard is under investigation for possible insider trading including claims CEO Bobby Kotick tipped off some investors to buy more shares before the $68.7bn Microsoft acquisition deal was announced. The American games maker, known for top series such as World of Warcraft and Call of Duty, said it was cooperating with the Securities Exchange Commission and the Department of Justice, according to a securities filing. "Activision Blizzard received a voluntary request for information from the SEC and a grand jury subpoena from the DOJ, both of which appear to relate to their respective investigations into trading by third parties - including persons known to Activision Blizzard's CEO - in securities prior to the announcement of the proposed transaction," the company stated on Friday in an 8-K submission. Continue reading * UK Prime Minister, Catalan groups 'targeted by NSO Pegasus spyware' UAE reportedly using 'legal' malware on erstwhile allies Thomas Claburn in San Francisco Mon 18 Apr 2022 // 20:17 UTC comment bubble on black Citizen Lab has reported finding suspected surveillance software on devices associated with both the UK Prime Minister's Office and what was formerly called the British Foreign and Commonwealth Office. The Canadian research outfit also said it had identified at least 65 individuals linked with Catalan civil society groups in Spain who were targeted by, or infected with, surveillance software. Catalonia is an autonomous region within Spain where there's an ongoing politically divisive fight for national independence. On Monday, Citizen Lab, a part of at the University of Toronto's Munk School, said it had found likely NSO Group Pegasus spyware infections on devices associated with UK Prime Minister Boris Johnson's office, 10 Downing Street, and on devices linked to the FCO, now called the FCDO, or the Foreign Commonwealth and Development office. Continue reading * TSMC's 2025 timeline for 2nm chips suggests Intel gaining steam Semiconductor world veteran says x86 titan could catch up with Asia-Pacific rivals in three years Dylan Martin Mon 18 Apr 2022 // 18:49 UTC 1 comment bubble on white TSMC said it won't start production at its 2nm node until the second half of 2025 or possibly the end of that year, which could signal a shift in the competitive landscape. The Taiwanese chip foundry revealed the timeline for its 2nm node, known officially as N2, during a conference call [PDF] last week for its first-quarter financial results. With a mid- to late-2025 production timeline, after late-2024 risk production, TSMC's 2nm production dies will likely land in the hands of their designers in volume in 2026, which, in turn, means those chips could, at the earliest, be available for phones, PCs, and servers that year. TSMC made the disclosure only a few days after Intel, which is revitalizing its competing foundry business, revealed that its next-generation 18A node will be ready for manufacturing in the second half of 2024, months ahead of the previously given 2025 timeline. As the A is short for angstroms, Intel's 18A label suggests it will be a 1.8nm process (see Register passim for caveats about node sizes.) Continue reading * Microsoft ups bug bounties 30% for cloud lines, pays more for 'scenario-based' exploits Plus: HP fixes critical Teradici flaws, Karakurt may be a Conti side hustle, and info-stealing malware set free Jessica Lyons Hardcastle Mon 18 Apr 2022 // 18:12 UTC comment bubble on black In Brief Microsoft will pay more -- up to $26,000 more -- for "high-impact" bugs in its Office 365 products via its bug bounty program. The new "scenario-based" payouts to the Dynamics 365 and Power Platform Bounty Program and M365 Bounty Program aim to incentivize bug hunters to focus on finding vulnerabilities with "the highest potential impact on customer privacy and security," Microsoft said late last week. Awards will increase as much as 30 percent in some cases, according to the Redmond software goliath. Continue reading * AI models to detect how you're feeling in sales calls Plus: Driverless Cruise car gets pulled over by police, and more Katyanna Quach Mon 18 Apr 2022 // 09:52 UTC 26 comment bubble on white In brief AI software is being offered to sales teams to analyze whether potential customers appear interested during virtual meetings. Sentiment analysis is often used in machine-learning research to detect emotions in underlying text or video, and the technology is now being applied to help people see how possible future clients are feeling in sales pitches to improve results, Protocol reported this month. The COVID-19 pandemic has moved a lot of meetings virtually as employees work from home. "It's very hard to build rapport in a relationship in that type of environment," said Tim Harris, director of product marketing at Uniphore, a software company specializing in conversational analytics. Continue reading * An early crack at network management with an unfortunate logfile It's a backronym, right? Richard Speed Mon 18 Apr 2022 // 07:30 UTC 34 comment bubble on white Who, Me? Come with us on a journey back to the glory days of Visual Basic 6, misplaced enthusiasm and an unfortunate naming incident. Welcome to Who, Me? Today's tale comes from a reader Regomised as "Stephen", who was working in the IT department of a Royal Air Force base. "My duties were many," he told us, "from running daily backups of an ancient engineering system using (I kid you not) reel-to-reel tapes to swapping out misbehaving printers." This being the early 2000s, his boss loaded up our hero with more tasks. He could change printers and tapes, so Visual Basic (and its bedfellow, Access) should present no problem. Continue reading * How to democratize ML? More public data, says MLCommons Foundation makes 30k hours of speech and 340k keywords in 50 languages available online Brandon Vigliarolo Sun 17 Apr 2022 // 09:43 UTC 8 comment bubble on white Unless you're an English speaker, and one with as neutral an American accent as possible, you've probably butted heads with a digital assistant that couldn't understand you. With any luck, a couple of open-source datasets from MLCommons could help future systems grok your voice. The two datasets, which were made generally available in December, are the People's Speech Dataset (PSD), a 30,000-hour database of spontaneous English speech; and the Multilingual Spoken Words Corpus (MSWC), a dataset of some 340,000 keywords in 50 languages. By making both datasets publicly available under CC-BY and CC-BY-SA licenses, MLCommons hopes to democratize machine learning - that is to say, make it available to everyone - and help push the industry toward data-centric AI. Continue reading * TACC Frontera's 2022: Academic supercomputer to run intriguing experiments Plus: Director reveals 10 million node hours, 50-70 million core hours went into COVID-19 research Brandon Vigliarolo Sat 16 Apr 2022 // 14:36 UTC comment bubble on black The largest academic supercomputer in the world has a busy year ahead of it, with researchers from 45 institutions across 22 states being awarded time for its coming operational run. Frontera, which resides at the University of Texas at Austin's Texas Advanced Computing Center (TACC), said it has allocated time for 58 experiments through its Large Resource Allocation Committee (LRAC), which handles the largest proposals. To qualify for an LRAC grant, proposals must be able to justify effective use of a minimum of 250,000 node hours and show that they wouldn't be able to do the research otherwise. Two additional grant types are available for smaller projects as well, but LRAC projects utilize the majority of Frontera's nodes: An estimated 83% of Frontera's 2022-23 workload will be LRAC projects. Continue reading * When the expert speaker at an NFT tech panel goes rogue Stick to the script, man! It's confusing enough already Alistair Dabbs Sat 16 Apr 2022 // 10:30 UTC 106 comment bubble on white Something for the Weekend How can you save the world's oceans? By investing in NFTs of course! A global network of campaigning filmmakers, Ocean Collective, hopes to drive up awareness about declining marine biodiversity by developing a digital Museum of Extinction. Items of artwork from the museum will then be sold as NFT purchases to raise cash to fund a documentary series on the topic along with other environmental awareness projects. Continue reading * Apple dev logs suggest 'nine new M2-powered Macs' 'Widespread internal testing' of four processor types Katyanna Quach Sat 16 Apr 2022 // 07:53 UTC 28 comment bubble on white Apple is seemingly testing four next-generation M2 processors on software developed by third-party app makers in at least nine Mac models that are likely to be upcoming laptops and desktops. Two years ago, the iGiant debuted its homegrown Arm-compatible M1 processor to power computers and iPads; the shift marked a departure from using x86 Intel silicon for its PCs. Instead of purchasing off-the-shelf processors, Apple - which was already designing its own mobile system-on-chips - wanted a custom design for its macOS products. Now it appears the M1's successor, the M2, is edging closer to launch, judging from developer logs leaked to Bloomberg that signal there is "widespread internal testing" of the chip family at Apple. Continue reading * Twitter preps poison pill to preclude Elon Musk's purchase plan Populist provocateur ponders partners to pay for platform prize Thomas Claburn in San Francisco Sat 16 Apr 2022 // 01:14 UTC 102 comment bubble on white Comment Twitter on Friday said its board of directors had unanimously approved a plan to prevent a hostile takeover, something that became a distinct possibility after billionaire Elon Musk offered $43 billion to buy the social media network. The poison pill, or "Rights Plan," the biz said, "will reduce the likelihood that any entity, person or group gains control of Twitter through open market accumulation without paying all shareholders an appropriate control premium or without providing the Board sufficient time to make informed judgments and take actions that are in the best interests of shareholders." The "Rights Plan" would require Musk to negotiate directly with the board to increase his share of the company beyond 15 percent. After that every existing shareholder, with the exception of Musk, would be able to buy Twitter stock at a discounted rate. Continue reading ABOUT US* * Who we are * Under the hood * Contact us * Advertise with us MORE CONTENT* * Latest News * Popular Stories * Forums * Whitepapers * Webinars SITUATION PUBLISHING* * The Next Platform * DevClass * Blocks and Files * Continuous Lifecycle London * M-cubed Situation Publishing The Register - Independent news and views for the tech community. Part of Situation Publishing SIGN UP TO OUR DAILY NEWSLETTER Subscribe Twitter Facebook LinkedIn feeds Security Scorecard no-js Biting the hand that feeds IT (c) 1998-2022 Do not sell my personal information Cookies Privacy Ts&Cs