[HN Gopher] Google has banned the training of deepfakes in Colab
___________________________________________________________________
Google has banned the training of deepfakes in Colab
Author : Hard_Space
Score : 179 points
Date : 2022-05-28 08:21 UTC (14 hours ago)
(HTM) web link (www.unite.ai)
(TXT) w3m dump (www.unite.ai)
| jakobov wrote:
| The regulation of ai begins
| thegreatdukd wrote:
| One thing I don't like is that they also banned SSH-ing to Colab,
| which means I can no longer remote SSH to VS Code and use Copilot
| on the stuff I am working on.
| dreamyfigment wrote:
| I wonder if we'll end up in a future where gaming GPUs will be
| restricted from running ML algorithms, to prevent private
| individuals from running advanced deepfakes/DALLE-10 or whatever.
| UmbertoNoEco wrote:
| That's a scary and realistic future.
|
| Any good (fiction or non-fiction) book on this? By "this" I
| mean governments trying to control overtly or secretly the use
| of computing resources. Bonus points if it is not Charlie
| Stross (I can't stand his writings).
| merlinscholz wrote:
| I mean Nvidia has already tried to do that with crypto mining
| Method-X wrote:
| This is terrifying. If only a certain class of people has
| access to this technology (i.e. governments and huge tech
| companies), the temptation to use that technology to control
| the class of people who don't will be too much.
| ahefner wrote:
| More likely the high end ML accelerators become so specialized
| as to diverge from GPUs. Carry on trying to use your gaming GPU
| for ML, with the knowledge that professionals are using
| hardware 100x as powerful.
| est31 wrote:
| There is this talk about the civil war on general purpose
| computing by Cory Doctorow:
| https://www.youtube.com/watch?v=HUEvRyemKSg
|
| He gave the same talk at Google:
| https://www.youtube.com/watch?v=gbYXBJOFgeI
|
| It's slowly becoming reality.
| alaricus wrote:
| Graphics is increasingly using ML, so I doubt it. But still,
| don't give them any ideas :)
| whatever1 wrote:
| Lol perfect timing for AWS https://studiolab.sagemaker.aws/
| orlp wrote:
| The results of GPT-3, Imagen, Github Copilot, etc, have shown me
| two things:
|
| 1. A critical part of AI that was missing no longer is: the model
| of our world. In natural language processing (more generally
| anything that interacts with humans) it has been an eternal issue
| that to understand sentences you must know the content, and
| therefore must, in theory, know 'everything' there is to our
| world and culture. As far as I'm concerned, the current large
| transformer models have captured this. There is genuine
| intelligence here, a lack of reasoning still but an abundance of
| knowledge.
|
| 2. We are entering a terrifying era of computing. The amount of
| resources needed to get the genuine AI above, are monstrous (at
| least with our current methods). Not only raw computing power,
| but also data. This means an unavoidable centralization and
| monopolization of something that I predict will be an essential
| component for human-computer interaction in the decades to come.
|
| What we're seeing here is a sign of (2).
|
| My prediction? Google and Facebook are here to stay. Their
| primary income right now is advertisement, I expect this to
| change over the next decade. Specifically these two companies are
| sitting on the biggest 'oil field' in modern history: an endless
| pipeline of modern human culture, ready to feed into ever larger
| models of our world.
| sinenomine wrote:
| While I partially agree with the theses, one should note that
| in absolute terms the compute spent even on the largest models
| of their kind is still pretty tame, compared to SV software
| engineer salaries: you could train your own Imagen for
| ~200000$. This is pretty much kickstarter-tier, and not at the
| higher end of the distribution.
|
| The web-scale data is also available: we already have a dataset
| superior to one Imagen used: https://laion.ai/laion-5b-a-new-
| era-of-open-large-scale-mult...
|
| Instead of despairing, we should train capable models in the
| open, like BigScience does:
| https://bigscience.huggingface.co/blog/what-language-model-t...
| and petition our governments to train & make available large
| models under permissive licenses, for the benefit of the public
| and small businesses.
|
| The big tech exceptionalism is overblown, you don't really need
| exotic engineering to train an LLM with Jax and a TPU pod, or
| whatever accelerator you managed to procure.
| orlp wrote:
| > you could train your own Imagen for ~200000$
|
| If my experience in machine learning (which admittedly is
| small scale) is applicable to these large-scale projects, I'd
| say that for each press-ready model there are probably dozens
| of previous attempts/versions that didn't make the cut. So
| I'd always add an order of magnitude on top of the final
| model cost.
|
| > The big tech exceptionalism is overblown, you don't really
| need exotic engineering to train an LLM with Jax and a TPU
| pod, or whatever accelerator you managed to procure.
|
| While I agree with you that the current best models are still
| within reach of motivated small-medium sized groups outside
| of big tech, if the scaling hypothesis holds up, the next big
| step forward is likely even larger and more expensive models.
| I was skeptical of the scaling hypothesis before, but much
| less so now.
|
| > petition our governments to train & make available large
| models under permissive licenses
|
| I would say that if it comes to this we are already firmly in
| the "centralization and monopolization" territory if we have
| to ask nation-states for their computational generosity.
| mupuff1234 wrote:
| The article and headline are pretty misleading imo, the ban seems
| like it's only for the free tier.
|
| Edit: Actually, it's not really clear.
| Hard_Space wrote:
| At the bottom of the FAQ section[1] that bans deepfakes, it
| says ' _Additional_ restrictions exist for paid users ', not
| 'different restrictions'.
|
| [1]
| https://research.google.com/colaboratory/faq.html#limitation...
| rockemsockem wrote:
| But if you click that link for "Additional Restrictions"
| there's nothing about deepfakes mentioned for pro.
|
| 3. Restrictions. In this Section 3, the phrase "you will not"
| means "you will not, will not attempt, and will not permit a
| third party to".
|
| When you use the Paid Service, you will not:
|
| share your Google account password with someone else to allow
| them to access any Paid Service that the person did not
| order; access the Paid Service other than by means authorized
| by Google; copy, sell, rent, or sublicense the Paid Services
| to any third party; use the Paid Service as a general file-
| hosting or media-serving platform; engage in peer-to-peer
| file-sharing; or mine cryptocurrency.
| machinekob wrote:
| Good for all of us maybe some good alternatives to colab will be
| created cause of that.
| JohnHaugeland wrote:
| why would any other company have the urge to do this?
| rslyr633 wrote:
| Because nature abhors a vacuum, and this technology has
| created one.
| JohnHaugeland wrote:
| It turns out that nature doesn't make SAAS
|
| This is an _extremely_ expensive service to make and run,
| requiring tens of millions of dollars of hardware to get
| started, and doesn 't produce revenue
|
| Instead of folksy sayings, see if you can answer the actual
| question
|
| What motivation does another company have for creating this
| service?
| rslyr633 wrote:
| >What motivation does another company have for creating
| this service?
|
| They won't need to create this deepfake training
| specifically, they just need to turn a blind eye to their
| existing hardware being applied this way.
|
| Three motivations come to mind:
|
| 1. It's a convenient selling point to grow the user base.
| "We give you complete freedom to run your models and
| don't spy on you like the other guys do! We will never
| stop you from doing things unless the government forces
| us to!"
|
| 2. It's an easy way to frontrun criminal activity that is
| poised to disrupt society in the near future. Hire a CEO
| with a cool sounding hacker name, advertise on a bunch of
| telegram groups, while secretly logging all identifiers
| and making all inputs/outputs searchable by law
| enforcement :)
|
| 3. By owning the data used to train and generate
| Deepfakes, you have a rich dataset that could be used
| academically or for training counter-models to detect
| future Deepfakes.
| JohnHaugeland wrote:
| I feel like you're completely missing the point of what I
| said
|
| 1. Not really
|
| 2. What?
|
| 3. No you don't
| loo wrote:
| > turns out that nature doesn't make SAAS
|
| E pur si existit.
| JohnHaugeland wrote:
| Quidquid Latin dictum sit does not, it turns out, altum
| sonatur
| machinekob wrote:
| Make people depends on ecosystem. I'll see
| AWS/Azure/NVIDIA doing smth like that as that just to
| create they own ecosystem as its extremely profitable to
| sell gpu VM's for them.
| JohnHaugeland wrote:
| Have you ever used colab? It's free
| visarga wrote:
| I presume it's like Microsoft turning a blind eye to
| unlicensed use of Windows and Office for home users. When
| they get hired they already prefer MS solutions they know
| since they were students.
| humanistbot wrote:
| They already exist. The market for SAAS Jupyter is already
| crowded. Azure Notebooks, Github Codespaces, Kaggle Kernels,
| AWS SageMaker, CoCalc, Datalore, Binder...
| machinekob wrote:
| Not even one is close to colab yet (first of google TPU is
| only way to access GPT-like models for most people, second
| you can get pretty neat GPU for few dollars when on other
| platforms you are paying a ton of money for that)
| sweetbitter wrote:
| If it's something that you think can be stopped, go ahead and try
| to stop it. Just know that you will be Open Sourced- even
| artificial general intelligence, when someday invented, will
| eventually have its design released in the clear.
| nestorD wrote:
| I _really_ do not like that because the ban feels fuzzy. If you
| are a researcher exploring those algorithms and using Colab
| because you have very limited GPU access of you own, you might
| now be worried that your very own topic of research will get you
| banned from the ressource you are using.
| ethbr0 wrote:
| I'd assume they recognized this, which is why it's implemented
| as a warning rather than a ban.
|
| Wouldn't be surprised if this is the first step towards
| correlating "Are you the 1% using huge amounts of Colab time to
| only run deepfake models?"
|
| We'll see what the follow-up is, but at least the initial
| implementation seems less auto-ban-y than more mature Google
| services default to.
| hugh-avherald wrote:
| Your complaint appears to be "funding is a constraint on
| research".
| chroem- wrote:
| By default, Google will not allow you to access GPU instances
| on regular GCP. The GPU quota is initialized to zero and you
| have to open a support ticket in order to request GPU access,
| justifying why your use case deserves access. Evidently,
| "personal research project" is not a valid use case, so that
| leaves Colab as the only other option in the Google
| ecosystem.
|
| AWS has a similar quota policy, but they seem less strict
| about use cases.
| peytoncasper wrote:
| I'm not so sure this is true. I was able to get access to a
| GPU quota of 1 on GCP within 2 minutes. I highly doubt
| anyone reviewed what I wrote. I'm not sure how many you
| were requesting access to, but I suspect the main reason
| for gating higher counts are cost related. Quite obviously
| they are not cheap.
| whymauri wrote:
| I've found that it varies. When I was at MIT, students
| and affiliates got credits and instances like candy (I
| feel like the school brand lowered friction). But then on
| Reddit, Stack Overflow, or Github I see a lot of people
| struggling to get quota (even when paying for the highest
| Colab tier).
| humanistbot wrote:
| It's only banned for the free tier.
| rockemsockem wrote:
| This small fact makes a huge difference. A researcher can
| certainly afford colab's modestly priced pro tier.
| alaricus wrote:
| How would such a ban even be enforced? This makes no sense?
| RandomThrow321 wrote:
| It probably just gives them the grounds to justify a ban, if
| necessary.
| endisneigh wrote:
| _free_ colab. You can do it if you pay.
| version_five wrote:
| I have wondered in the past what kind of monitoring / data
| harvesting google is doing with colab. I don't use them for
| anything nontrivial or work related precisely because I worry
| about google stealing my stuff or that I would be violating
| agreements with my customers by exposing their data to google.
| This story reinforces that they are snooping.
|
| That aside, I don't think I find this too upsetting. They don't
| let people e.g. mine cryptocurrency either, I think it's fine
| that they limit how you can use their free (or cheap) "service".
| summerlight wrote:
| > I worry about google stealing my stuff or that I would be
| violating agreements with my customers by exposing their data
| to google.
|
| Perhaps it just follows Google cloud's user data policy as long
| as you're using a paid service? But the latter might be a
| legitimate concern; some customer may ask their data not to
| leave any networks outside the contract and depending on this
| types of public services will increases the likelihood of such
| accidents.
| ClassyJacket wrote:
| Is there a good alternative to Colab you can just pay for,
| that's good for hobbyist individuals? Especially one that gives
| you shell access?
|
| I tried using PaperSpace for a while but it works maybe 1 in 3
| times I tried to use it, they (temporarily) banned me for no
| reason, they send me "itemized" bills with empty line items,
| and their support was no help.
| kmeisthax wrote:
| JupyterLab
|
| I've migrated several notebooks from Colab to JupyterLab when
| playing around with various generative art models; it's a
| drop-in replacement and runs on your local machine's
| hardware.
|
| I _did_ have to do some fiddling to install a bunch of Nvidia
| crap on my local machine, however; blame CUDA.
| nmiculinic wrote:
| Disclaimer: I work there.
|
| https://www.grid.ai/
| version_five wrote:
| Yes, I like them (and I feel like I've seen some other
| similar services but haven't looked into them closely). Its
| easy to start and stop interactive instances (easier than
| AWS imo) for doing stuff in notebooks, and to run training
| jobs on more powerful GPUs when you need to, because you're
| only paying for it when you're actually using it.
|
| OVH actually has a (distantly) similar service (much less
| user friendly than grid and without the whole "grid"
| feature, but that matters less for hobbyists if you're not
| doing a lot of parallel runs). I liked the OVH one a lot in
| principle, but in practice found it too buggy to use
| properly (and they don't have customer support). For a
| budget project it could be worth trying.
| endisneigh wrote:
| What's wrong with Colab Pro out of curiosity (assuming you
| don't already know of it, which isn't restricted like the
| free version)?
| ClassyJacket wrote:
| Last time I checked it wasn't available outside the US and
| Canada at all, and didn't have a terminal, but I just
| checked and both of those seem to have changed.
|
| I'm also not a fan of having to upload files to Google
| Drive and then write special code to import them. It means
| my code isn't portable to run locally or anywhere else and
| it's just extra effort. I just want normal file access and
| to write scripts like I normally would.
|
| Still, I was spurred to ask the question because some
| people are categorically opposed to anything Google due to
| privacy issues.
| minimaxir wrote:
| Technically that's not the target use case for Colab;
| it's Notebook in, Notebook out.
|
| That's a use case more suited to running a VM itself.
| version_five wrote:
| What I settled on is I have a computer with a lower-end
| nvidia gpu that lets me debug code and do basic stuff but is
| impractical for training any large model or with lots of data
| (although since the incremental cost is only the electricity,
| you can often still get far by training overnight or when
| you're doing something else). And then I use AWS more
| deliberately once I've figured out exactly what I want to
| run. Otherwise, you're reserving a cloud GPU for all the
| setup and debugging.
|
| If getting a gpu is not practical, I'd try one of the K80s or
| other cheaper GPUs on AWS for all your debugging and then
| port to the more expensive instances only if you need them.
| Doing this as a hobbyist will be relatively inexpensive. I
| think a lot of the cost incurred as a hobbyist comes from
| reserving a gpu when you don't actually need it.
| jph00 wrote:
| PaperSpace has improved a lot in the last few months. I've
| been using them nearly exclusively recently and have had zero
| issues.
|
| They provide persistent storage and access to the full
| JupyterLab or Jupyter Notebook software, unlike Colab. I find
| this makes life far easier, since all my normal terminal
| workflows work fine.
| merlinscholz wrote:
| Azure and JetBrains both offer jupyter notebooks too
| fxtentacle wrote:
| You could use my Colab imitation docker image on OVH:
| https://hub.docker.com/r/fxtentacle/ovh-colab-sagemaker-
| comp...
|
| OVH charges $2 per hour for a V100S
| krinn_silver wrote:
| https://deepnote.com is exactly this!
|
| Disclaimer: I work for Deepnote
| thorum wrote:
| Lambda is great, if you can afford it:
| https://lambdalabs.com/service/gpu-cloud/pricing
| albert_e wrote:
| - Amazon Sagemaker Studio Lab (free, email signup, no AWS
| account or credit card required)
|
| paid services on AWS:
|
| - Amazon Sagemaker Studio
|
| - Amazon Sagemaker Notebook Instances (think EC2 + jupyter +
| integration with AWS services)
| jph00 wrote:
| Studio Lab doesn't have any GPU instances available in
| practice, unfortunately. They're permanently out of
| resources at the moment. The other two options are pretty
| expensive.
| [deleted]
| fartcannon wrote:
| They do all the harvesting. Everything they can. Every morsel
| of data they can get, they get.
|
| Same thing Microsoft does with Windows, github and VSCode.
|
| Just hoovering up all the ideas you have.
| endisneigh wrote:
| if you can prove this you can make a lot of money
| Terry_Roll wrote:
| Get real, you obviously dont know how organised crime
| works.
| endisneigh wrote:
| You have proof of these incredible claims - seems crazy
| to think Google/Microsoft are stealing private data? I
| know some folks at the NYT if you can provide credible
| evidence.
| Terry_Roll wrote:
| Nah the house has been entered without permission hard
| drives wiped, been shot at, few attempts on my life,
| threatened countless times, you obviously live in lala
| land! You dont run a country by being nice you know!
| fartcannon wrote:
| Who do you know at NYT?
| fartcannon wrote:
| Oh yeah, how?
| endisneigh wrote:
| Class action, violation of the Terms and Conditions.
|
| I'd honestly be shocked if Google/Microsoft are
| monitoring the clients of their public cloud and keeping
| the data to create their own products off of.
| fartcannon wrote:
| Prepare to be shocked.
| endisneigh wrote:
| proof?
| fartcannon wrote:
| The existence of copilot, for one.
| endisneigh wrote:
| that's not proof, lol. That uses publicly available code.
| are you suggesting it's trained on private code? not to
| mention I was talking about cloud workloads since you
| said _everything_
| fartcannon wrote:
| Nobody said anything about private or public. What was
| said was 'everything they can get their hands on'. Thou
| doth protest too much.
| loo wrote:
| That includes private data. Thou critical thinkest too
| little.
| schoen wrote:
| I think you want "thinkest" (2nd person) rather than
| "thinketh" (3rd person).
|
| https://en.wiktionary.org/wiki/thinkest
| loo wrote:
| Thanks. Edited.
| endisneigh wrote:
| everything they can get their hands on includes both
| private and public data?
|
| they have access to the private data, by definition. the
| entire point is that it's not exposed nor is it used.
| clearly you've never worked at one of the companies in
| question lol. when you're on-call you might need to
| actually access private data.
|
| true data sovereignty products are something countries in
| the EU are wanting USA companies to implement though, so
| then it's simply inaccessible, not public or private.
|
| the industry is moving in this direction with things like
| AWS Outposts
| fartcannon wrote:
| Google needs to spend more on PR I think.
| filoleg wrote:
| > _when you 're on-call you might need to actually access
| private data._
|
| When I was on-call and had to dive deep to solve
| customers' issues at one of those companies, we could
| only access anonymized data stripped of all PII
| (including emails and even IP addresses), to the point
| where the only non-scrambled things we could get were
| timestamps and error stacktraces/executed commands (which
| were also stripped of all PII, which made debugging a bit
| more difficult).
|
| And every instance of access to that anonymized no-PII
| data required a physical authenticator, and it would
| automatically send an email to your manager (and some
| mailing group, i forgot) and create a special log that
| indicated "such and such was trying to access this
| anonymized ABC data for XYZ stated purpose".
| willcipriano wrote:
| The terms and conditions almost certainly bind you to
| arbitration so no class action.
| andrewnicolalde wrote:
| What about Europe?
| Datenstrom wrote:
| I'd be shocked if they weren't, it seems like a totally
| standard big corp playbook move. It would be the same
| exact thing Amazon was doing with their independent
| seller data[1].
|
| [1]: https://www.wsj.com/articles/amazon-scooped-up-data-
| from-its...
| endisneigh wrote:
| that's not the same thing, since it's public info.
| [deleted]
| Spooky23 wrote:
| They'd need to be very careful in doing so, as they both
| operate services that are certified and audited to meet
| US Gov FedRAMP High control standards, and doing that
| would violate several audited control standards.
|
| A breach in that area would cost them potentially
| billions in losses.
| [deleted]
| onesmalldrop wrote:
| Public clouds are separated from their gov versions, with
| gov versions using older, modified versions of what
| everyone else uses. public clouds ARE NOT fedramp
| compliant
| rrdharan wrote:
| This is not true. Google Cloud 's FedRAMP offering uses
| the same regions, zones and binaries and configurations
| as the regular public cloud offering.
|
| Source: I work on FedRAMP compliance for GCP Databases.
| Spooky23 wrote:
| Not necessarily. Lots and lots of .gov operates in
| commercial clouds. Google doesn't even have a distinct
| offering - just support add-ons for things like CJIS
| compliance. Microsoft's footprint is a lot more complex.
|
| Outside of the DoD space, the main distinctions between
| these offerings is where data can be stored and where and
| to what level of vetting vendor employees are.
| amusedcyclist wrote:
| Your ideas aren't worth that much to be honest
| malcolmgreaves wrote:
| Good! Deepfakes are going to be a massive problem for society
| going forward. Glad to see Google is at least recognizing it's
| responsibility as a compute platform here.
| rockemsockem wrote:
| It's only restricted on the free tier.
|
| Honestly they're probably finding that, like cryptocurrency
| miners, free tier deepfake users born burn lots of GPU time
| compared to others who might run a couple GPU tasks for a
| little while to get a few results or test things.
| mark_l_watson wrote:
| It is their right to ban any use they want. Colab is an awesome
| service, so much so that I seldom use my at-home GPU rig.
|
| As consumers we can use or not use platforms like Google,
| Twitter, etc. as we like, or not.
|
| Colab's ban will be good for small 3rd party GPU cloud providers
| and that is a good thing. Same thing with Twitter: in the last
| month there is now better material on Mastodon.
| DarylZero wrote:
| edmundsauto wrote:
| This is not an appropriate response - it does not advance the
| conversation, and is clearly derogatory towards OP. Try to do
| better next time.
| nmilo wrote:
| Yes, that's how society works. If you use a service
| respectfully you get rewarded with the ability to keep using
| that service.
| DarylZero wrote:
| Subservient lapdog attitude toward corporations is a modern
| thing not a generic property of "society."
| passivate wrote:
| And the authors are using their rights to free expression to
| criticize Google. No rights have been taken away. So its all
| good.
| visarga wrote:
| > Colab is an awesome service, so much so that I seldom use my
| at-home GPU rig.
|
| Can you use Colab from VSCode? Notebooks are great for
| exploration but not for complex projects.
| davidatbu wrote:
| So I really dislike Jupyter, and I've tried using this[0]
| before to ssh into Colab and do work in a terminal setup.
|
| You have to be careful to back up your code frequently (what
| I did was push to Github) since your ssh session goes away
| when your "kernel" (or whatever your colab session is called)
| goes away. You wouldn't have this worry if you were just
| using the web interface, since the code is always saved.
|
| I'd imagine that, if you want to use VSCode, it's remote
| editing features, which I keep on hearing a lot about, would
| come in handy to edit over ssh!
|
| [0] https://github.com/WassimBenzarti/colab-ssh
___________________________________________________________________
(page generated 2022-05-28 23:01 UTC)