https://scottlocklin.wordpress.com/2025/10/31/stochastic-computing/ Skip to content Skip to search - Accesskey = s Locklin on science Stochastic computing Posted in non-standard computer architectures by Scott Locklin on October 31, 2025 I've wanted to write about this topic since I started this blerg, but various things have kept me from it. Stochastic computing is a subject I've been somewhat aware of since I went through the " Advances in Computers" series in the LBNL library while procrastinating over finishing my PhD thesis. It's an old idea, and the fact that sci-hub exists now means I can look up some of these old papers for a little review. The fact that there is a startup reviving the idea (their origin paper here) gives a little impetus to do so. Von Neuman had a vague idea along the lines of stochastic computing in the 1950s, but the first actual stochastic computing architectures were done more or less independently by Brian Gaines, and Wolfgang (Ted) Poppelbaum in the mid-1960s. I first read about this in "Advances in Computers" in the LBNL library while procrastinating on writing my dissertation, and while I didn't take any notes, it always remained stuck in my noggin. The once-yearly journal started in 1960 and I at least skimmed the abstracts all the way to 2002 or whenever it was I was procrastinating (physics libraries are way better than the internet/wikipedia for this sort of thing). My takeaway was that Tukey's FFT was godawful important, there were a few cool alternatives to normal computer architectures, and climatological calculations have always been more or less the same thing (aka lame) with bigger voxels in the old days. Stochastic computing was one of the cool alternatives; it was competitive with supercomputers back in the 1960s, and it was something a small research group could physically build. As a reminder, a $25 million (current year dollars) CDC6600 supercomputer of 1964 was two megaflops with space for 970 kilobytes of RAM. They made about 100 of them total for the world. An NVIDIA titan RTX has 24gb and runs at something like 130 teraflops. Your phone is probably 2 teraflops. Something like stochastic gradient descent was, therefore, not worth thinking about even on a supercomputer in those days. Fitting a perceptron or a Hidden Markov model was a considerable technical achievement. People would therefore use various kinds of analog computer for early machine learning research. That includes stuff involving matrices which were modeled by a matrix of resistors. Analog computers can do matrix math in O(1); it only takes as long as the voltage measurement takes. Analog used to be considered parallel computation for this reason. One can think of stochastic computing as a way of digitally doing a sort of analog computing. Instead of continuous voltage levels, stochastic computing used streams of random (known distribution) bits and counted them. This has many advantages; counting bits is pretty easy to do reliably, where measuring precise voltages and reliably converting them back into bits requires much more hardware. It's also more noise immune in that a noise event is an extra bit or two which doesn't affect the calculation as much as some unknown voltage would affect an analog machine. Thinking about this a different way, if you're dealing with machine learning, the numbers being piped around represent probabilities. A bit of noise hitting a stochastic stream of bits might add a couple of bits and is unlikely to make the overall probability greater than 1. If your probability is analog voltage levels, any old noise is very likely to botch your probability and turn it into some ad-hoc Dempster-Shafer thing (which hadn't been invented yet) where probabilities sum to greater than 1. While I'm talking about probabilities here, let it be known that standard mathematical operations are defined in these computers: addition, multiplication (including for matrices), calculus: just like with analog computers. Noise immunity is quite a useful property for 1960s era computers, especially in noisy "embedded" environments. TTL was available, but noisy and not particularly reliable for digital calculations. It was contemplated that such machines would be useful in aircraft avionics, and it was actually used in a small early autonomous neural net robot, tested in a British version of LORAN, a radar tracker and PID system. The early stochastic computers were generally parts of some kind of early neural net thing. This was a reasonable thing to do as back then it was known that animal neural nets were pulse rate encoded. The idea lives on in the liquid state machine and other spiking neural nets, a little-loved neural net design which works sort of like actual neurons, but which is quite difficult to simulate on an ordinary computer that has to calculate serially. Piece of cake for stochastic computers, more or less, as stochastic computers include all the parts you need to build one, rather than simulating pulses flying around. Similarly, probabilistic graphical models are easier to do on this sort of architecture, as you're modeling probabilities in a straightforward way, rather than the indirect random sampling of data in used PGM packages. Brian Gaines puts the general machine learning problem extremely well, "the processing of a large amount of data with fairly low accuracy in a variable, experience-dependent way." This is something stochastic machines are well equipped to deal with. It's also a maxim repeatedly forgotten and remembered in the machine learning community: drop-out, ReLu, regularization, stuff like random forests are all rediscoveries of the concept. This isn't a review article, so I'm not mentioning a lot of the gritty details, but there are a lot of nice properties for these things. The bit streams you're operating on don't actually need to be that random: just uncorrelated, and there is standard circuitry for dealing with this. Lots of the circuits are extremely simple compared to their analogs on a normal sequential machine; memory is just a delay line for example. And you can fiddle with accuracy of digits simply by counting fewer bits. Remember all this 1960s stuff was done with individual transistors and early TTL ICs with a couple devices on a chip. It was considered useful back then and competitive with and a lot cheaper than ordinary computers. Now we can put down billions of far more reliable transistors down on a single chip. Probably more unreliable transistors, which is what stochastic computers were designed around. Modern chips also have absurdly fast and reliable serial channels, like the things that AMD chiplets use to talk to each other. You can squirt a lot of bits around, very quickly. Much more quickly than in the 1960s. Oh yeah and it's pretty straightforward to map deep neural nets onto this technology, and it was postulated back in 2017 to be much more energy efficient than what we're doing now. You can even reduce the voltage and deal with the noise byproduct much more gracefully than a regular digital computer. I assume this is what the Extropic boys are up to. I confess I looked at their website and one of the videos and I had no idea WTF they were talking about. Based Beff Jezos showing us some Josephson gate prototype that has nothing to do with anything (other than being a prototype: no idea why it had to be superconducting), babbling about quantum whatevers did nothing for me other than rustling my jimmies. Anyway, if their "thermodynamic computing" is stochastic computing, it is a very good idea. If it is something else, stochastic computing is still a very good idea and I finally wrote about it. They're saying impressive things like 10,000x energy efficiency, which of course is aspirational at the moment, but for the types of things neural nets are doing, it seems doable to me, and very much worth doing. In my ideal world, all the capital that has been allocated to horse shit quasi-scams like OpenAI and all the dumb nuclear power startups around it (to say nothing of the quantum computard startups which are all scams) would be dumped into stuff like this. At least we have a decent probability of getting better computards. It's different enough to be a real breakthrough if it works properly, and there are no obviously impossible lacunae as there are with quantum computards. Brian Gaines publications: https://gaines.library.uvic.ca/pubs/ Good recent book on the topic: Extropic patents: https://patents.google.com/?assignee=extropic&oq=extropic Share this: * Click to share on Facebook (Opens in new window) Facebook * Click to share on X (Opens in new window) X * Click to share on Reddit (Opens in new window) Reddit * Like Loading... Related 8 comments << Pre-Dreadnaughts: an aesthetic appreciation Things that should be considered essential vitamins but aren't >> 8 Responses Subscribe to comments with RSS. 1. Dennis's avatar Dennis said, on October 31, 2025 at 5:30 pm Hi Scott, I have to say I'm a bit surprised to see you enthusiastic about stochastic computers. Aren't those so-called quantum computing companies (like D-Wave's quantum annealers, but also the supposed gate-based quantum computers) in the end actually producing -- in effect -- such stochastic machines? I mean, apart from all the talk about a universal quantum computer, if what they are actually building is a stochastic computer, then at least something concrete has come out of it. (The 2025 arXiv paper you reference, for instance, mentions quantum annealers that naturally produce bitstrings stochastically and have weights to control the distribution in [30].) Reply + Scott Locklin's avatar Scott Locklin said, on October 31, 2025 at 5:42 pm D-wave's thing is bullshit woo. I assume they're mentioning quantum for the woo factor, which doesn't speak well for their marketing efforts (the fact that I didn't know WTF they were talking about after reading the paper and the marketing is another demerit on marketing). Their thing is made with transistors. Assuming it's the old stochastic computers it is a fine idea. If it's something else, who knows. Reply 2. ahgamut's avatar ahgamut said, on November 1, 2025 at 2:06 am I looked through the Gaines 1967 paper "Stochastic Computing" to see what's going on. It seems the idea there is to encode a number p on the unit interval as an infinite sequence of independent bits, each of which can be 1 with probability p. Ok, so multiplication via AND gates, addition with some accounting, and then I guess other arithmetic operations show up via cancellation rules. The paper mentions O(N^2) levels for each continuous variable. Throwing more bits across for accuracy makes sense (tighter confidence interval when estimating p). But generating so many independent bits appears a bit weak, as does the collection at the end. I assume there will be some custom PRNGs/hardware to handle both. Interesting idea! The constraints appear to be with storage/ conversion, but nothing stands out as a major dealbreaker yet, which is a bit surprising. Let's see. Reply + Scott Locklin's avatar Scott Locklin said, on November 1, 2025 at 12:14 pm Gaines historical chapter in the above mentioned 2019 book pointed out even back in the day they were better off punting on a regular computer than some new kind of weird analog machine: they'd find a use for it eventually. Despite his invention of the idea of stochastic computing, he spent a few years developing an early PDP-7 kind of thing he designed himself. Was easier to sell that kind of computer. Anyway it looks like a good idea to me. If nothing else, it could make for very energy efficient math coprocessors. I'll point out again I only have a vague idea if extropic is actually doing this through their impenetrable wall of "sciencese." Must be something in the ballpark though. Reply o ahgamut's avatar ahgamut said, on November 4, 2025 at 9:33 am The extropic thing showed up for me elsewhere, so I took a look at their blog. Seems like Gibbs sampling but with custom hardware? The energy-based stuff looks similar to handling the unnnormalized densities that show up when dealing with MCMC (ie take ratio/log for relative direction). PGMs are nice, custom hardware for efficiency is nice, but Gibbs/MCMC-type ideas have known issues (getting stuck in a mode, not able to explore the whole space, convergence criteria). Not sure what the workaround is (diffusion models?). The "pbit" stuff seems like the stochastic computing mentioned in your blog post, but not sure about that either. There's a python package, I'll see if that helps. Reply # Scott Locklin's avatar Scott Locklin said, on November 4, 2025 at 1:25 pm I think the solution to PGMs is same as it was for deep nets; throw more compute at it. If you click on their patent applications, which are a steaming pile of gibberish as far as I can tell, they're also into neural thingees. In particular, of course, contrastive divergence stuff (also Gibbs sampled). The pbit mentions and the fact that they call it "thermodynamic" are the only real connection I can make to old school stochastic computing. It's entirely possible it is something completely different, or is complete bullshit (their texts are not helpful; it's entirely in ridiculous "sciencese"), but I have wanted to talk about stochastic computing literally since I started the blog as it's little known and a really cool approach. If that's what they're doing, good on them. If not, shame on them for being such poor communicators. Actually shame on them for that anyway; it's really bad! They obviously spent more time on the aesthetics of their prototype than they did on telling people how it works. Reply 3. Anon's avatar Anon said, on November 1, 2025 at 1:38 pm I have a few things I want to drop in one comment. First, here's a book I stumbled upon from the 60s that shows you how to make a digital computer out of household items, like paperclips, light bulbs, empty thread spools, etc. https://archive.org/details/ howtobuildaworkingdigitalcomputer_jun67 I wonder how many IFLS soybois could conceive of something like this. Next, I think you'll enjoy these slides: https://cr.yp.to/talks/2015.04.16/slides-djb-20150416-a4.pdf He mentions LAPACK and OpenBLAS, and various reasons why compilers will basically always suck. Knuth is quoted at the end as saying that hand-crafted assembly will always be faster than the best optimizing compilers. It reminds me of your old posts about numeric linear algebra and all the different optimizations and bottlenecks to overcome. I especially remember when you talked about GotoBLAS, which is basically all the proof you need: Goto-san's hand-optimized assembly blew the best compiled code in the world out of the water, and it ran the fastest supercomputers in the world for years. Finally, something to make you laugh. I remember reading somewhere on here where you said "10 nm" ain't really 10 nm, and then upon a little further reading, I found this gem. https://www.eejournal.com/article/no-more-nanometers/ For many decades (up until 2016), the International Technology Roadmap for Semiconductors (ITRS) told us years in advance what each node should be named. ITRS is a committee that (best I can determine) had regular meetings where they took the previous node name, divided by the square root of two, rounded to the nearest integer, declared that result to be the new node name, and then drank lots of wine. Beginning in 2016, they re-named themselves "International Roadmap for Devices and Systems" - taking on a much broader system-level charter, and assuring that they'd have an excuse to drink wine well beyond the impending end of Moore's Law. Pretty much says it all. Reply + tcnymex's avatar tcnymex said, on November 2, 2025 at 8:48 am I had that book. Right after I built my three and a half foot tall Saturn V model, I built that computer. That was in the spring, that summer we moved to Westchester County into the same town where the chairman of IBM lived, back when IBM was bigger than the all the BUNCH companies combined (Burrows Univac NCR ControlData and Honeywell? I think you could have thrown Germany's Nixdorf, Italy's Olivetti, and whatever British General Electric computer company there was, and IBM would still have been bigger). When the chairman of IBM lives in your town, (and his kids go the public school,) you can bet there will be 3 IBM 2741 Selectric terminals at school (one for 6th, one for 7th, and one for 8th grade ... And how gauche that the kid from Brooklyn thought he was going to a "Junior high" instead of a "middle school") ... So right after I built my own working digital computer, I stopped using it in favor of APL on a terminal that typed back at me at a blazing 15 characters per second. (Which meant printing out an. ASCII art Snoopy took about 12 minutes.) By the end of the year I had figured out there was more to computing than printing out ASCII art Snoopys, and how cool it was to be in 7th grade running APL on an IBM Stretch 90, (there were only six Stretch-90s in the world, two of which were at Livermore, with a third at Los Alamos, and for some reason a fourth just outside of Washington at something called "the Maryland Purchasing Commission") Reply Leave a comment Cancel reply [ ] [ ] [ ] [ ] [ ] [ ] [ ] D[ ] About me: * Stuff I like * The Futurist Manifesto * About Scott Locklin Past blogs Past blogs[Select Category ] Email Subscription Enter your email address to subscribe to this blog and receive notifications of new posts by email. Email Address: [ ] Sign me up! Join 2,489 other subscribers RSS link thingee * RSS - Posts * RSS - Comments Create a free website or blog at WordPress.com. * Comment * Reblog * Subscribe Subscribed + [330189] Locklin on science Join 1,628 other subscribers [ ] Sign me up + Already have a WordPress.com account? Log in now. * + [330189] Locklin on science + Subscribe Subscribed + Sign up + Log in + Copy shortlink + Report this content + View post in Reader + Manage subscriptions + Collapse this bar %d [b]