[HN Gopher] Memory access on the Apple M1 processor
       ___________________________________________________________________
        
       Memory access on the Apple M1 processor
        
       Author : luigi23
       Score  : 359 points
       Date   : 2021-01-06 17:06 UTC (1 days ago)
        
 (HTM) web link (lemire.me)
 (TXT) w3m dump (lemire.me)
        
       | lrossi wrote:
       | Shouldn't you choose the random numbers such that array[idx1] ^
       | array[idx1 + 1] are guaranteed to fall in the same cache line?
       | Assuming that it has that. Right now some accesses cross the end
       | of the cache line.
        
         | CyberRabbi wrote:
         | Technically you are correct but it's expected to cross a cache
         | line 1/16 times (or however many ints there are in a cache
         | line). There is an implicit assumption that that is relatively
         | infrequent enough that it shouldn't increase the average time
         | too much, but that assumption should be tested.
        
       | jayd16 wrote:
       | >our naive random-memory model
       | 
       | Doesn't everyone use the (I believe) still valid concepts of
       | latency and bandwidth?
        
         | sroussey wrote:
         | Depends on context.
         | 
         | For example, what is the bandwidth and latency when you ask for
         | the value at the same memory address in an infinite loop? And
         | how does that compare to the latency and bandwidth of a memory
         | module you buy on NewEgg?
        
           | fluffy87 wrote:
           | L1 BW.
           | 
           | When people use BW in their performance models, they don't
           | use only 1 bandwidth, but whatever combination of bandwidth
           | makes sense for the _memory access pattern_.
           | 
           | So if you are always accessing the same word, the first acces
           | runs at DRAM BW, and subsequent ones at L1 BW, and any
           | meaningful performance model will take that into account.
        
         | whoisburbansky wrote:
         | The concepts are still broadly valid, the naivety being
         | referred to is the assumption that two non adjacent memory
         | reads will be twice as slow as one memory read or two adjacent
         | reads.
        
         | wyldfire wrote:
         | How do latency and bandwidth relate to the cost model for the
         | code in the benchmark?
         | 
         | When creating the model discussed in the post, we're using it
         | to try to make a static prediction about how the code will
         | execute.
         | 
         | Note that the goal of the post is not to merely measure the
         | memory access performance, it's to understand the specific
         | microarchitecture and how it might deliver the benefits that we
         | see in benchmarks.
        
       | foota wrote:
       | Is this per core or shared between cores?
        
         | hundchenkatze wrote:
         | Per core I think, emphasis is mine.
         | 
         | > It looks like a _single_ core has about 28 levels of memory
         | parallelism, and possibly more.
        
           | foota wrote:
           | I was wondering if this might be a shared resource though,
           | since it doesn't seem they tested with multiple threads.
        
       | wrsh07 wrote:
       | Ok, summary:
       | 
       | This article lays out three scenarios: 1) accessing two random
       | elements
       | 
       | 2) accessing 3 random elements
       | 
       | 3) accessing two pairs of adjacent elements (same as (1) but also
       | the elements after each random element)
       | 
       | It then does some trivial math to use the loaded data.
       | 
       | A naive model might only consider memory accesses and might
       | assume accessing an adjacent element is free.
       | 
       | On the Mac m1 core, this is not the case. While the naive model
       | might expect cases 1 & 3 to cost the same and case 2 to cost 50%
       | more, instead cases 2 & 3 are nearly the same (3 slightly faster)
       | and case 2 is about 50% more expensive than 1.
        
         | jayd16 wrote:
         | I don't really understand the comparison because it seems like
         | scenario 3 (2+) is doing more XORs and twice the accesses to
         | array over the same amount of iterations.
         | 
         | We have to assume these are byte arrays, yes? Or at least some
         | size that's smaller than the cache line. You would still pay
         | for the extra unaligned fetches. I don't think this is a valid
         | scenario at all, M1 or not.
         | 
         | Anyone want to run these tests on an Intel machine and let us
         | know if the authors "naive model" test hold there?
        
           | wrsh07 wrote:
           | The point of the naive model is that you assume memory
           | accesses dominate
           | 
           | That is, the math part is so trivial compared to the memory
           | access that you could do a bunch of math and you would still
           | only notice a change in the number of memory accesses.
           | 
           | Also it looks like the response to yours links their test and
           | the naive model predicts correctly
        
             | jayd16 wrote:
             | I think 5% is a non-trivial difference but alright, its a
             | much bigger difference on the M1.
             | 
             | I guess I still don't understand whats going on here.
             | 
             | Scenario 1 has two spatially close reads followed by two
             | dependent random access reads.
             | 
             | Scenario 3 (2+) has two spatially close reads, and two
             | pairs of dependent random access reads of two spatially
             | close locations.
             | 
             | Why does it follow that this is caused by a change in
             | memory access concurrency? The two required round trips
             | should dominate both on the M1 and an Intel but for some
             | reason the M1 performs worse than that. Why?
             | 
             | I can't help but feel the first snippet triggers some SIMD
             | path while the 3rd snippet fails to.
        
               | wrsh07 wrote:
               | I think the 5% can maybe be accounted for by the cache
               | line (you mentioned this above, and I don't think the
               | experiment does anything to prevent the issue)? If it's
               | 1/16th chance of crossing the cache line, that maybe is
               | about 5% of the time? I say that with pretty low
               | confidence though
               | 
               | I think you raise a good question, though -- what really
               | is going on here? Is this just a missed optimization
               | compiling for the m1?
               | 
               | Or is it actually something fundamental about how reads
               | happen with an m1? I'm definitely not knowledgeable
               | enough to know how to answer this
        
           | africanboy wrote:
           | I did it on an old i7 laptop
           | 
           | https://news.ycombinator.com/item?id=25661055
        
         | temac wrote:
         | > A naive model might only consider memory accesses and might
         | assume accessing an adjacent element is free.
         | 
         | Really depends on the level of naivety and the definition of
         | "free". It would be less insane to write that: accessing an
         | adjacent element has a negligible overhead if the data must be
         | loaded from RAM and there are some OOO bubbles to execute the
         | adjacent loads. If some data are in cache the free adjacent
         | load claim immediately is less probable. If the latency of a
         | single load is already filled by OOO, adding another one will
         | obviously have an impact. If the workload is highly regular you
         | _can_ get quite chaotic results when making even some trivial
         | changes (even sometimes when _aligning the .text differently_!)
         | 
         | And the proposed microbenchmark is way too simplistic: it is
         | possible that it saturates some units in some processors and
         | completely different units in others...
         | 
         | Is the impact of an extra adjacent load from RAM _likely_ to be
         | negligible in a real world workloads? Absolutely. With precise
         | characteristics depending on your exact model  / current freq /
         | other memory pressure at this time, etc.
        
       | willvarfar wrote:
       | A lot of commenters here are saying that Apples advantage is that
       | it can profile the real workloads and optimise for that.
       | 
       | Well that's true and could very well be an advantage. An
       | advantage in that they did it, not in that only they have access
       | to it.
       | 
       | Intel and AMD can trivially profile real world workloads too.
       | 
       | Did they? I don't know what Apple did, but the impression I get
       | is that intel certainly hasn't.
        
         | gameswithgo wrote:
         | of course every cpu designer is using real world workloads to
         | guide design
        
       | nabla9 wrote:
       | What is the cache line size and page table size in M1?
       | sysconf(_SC_PAGESIZE); /* posix */
       | 
       | Can you get direct processor information like LEVEL1_ICACHE_ASSOC
       | and LEVEL1_ICACHE_LINESIZE from the M1??
        
         | momothereal wrote:
         | `getconf PAGESIZE` returns 16384 on the base M1 MacBook Air.
         | 
         | The L1 cache values aren't there. The macOS `getconf` doesn't
         | support -a (listing all variables), so they may just be under a
         | different name.
         | 
         | edit: see replies for `sysctl -a` output
        
           | lilyball wrote:
           | Is it possibly exposed via sysctl, which does support a flag
           | to list all variables?
        
             | messe wrote:
             | From sysctl -a on my M1:
             | hw.cachelinesize: 128         hw.l1icachesize: 131072
             | hw.l1dcachesize: 65536         hw.l2cachesize: 4194304
             | 
             | EDIT: also, when run under Rosetta hw.cachelinesize is
             | halved:                   hw.cachelinesize: 64
             | hw.l1icachesize: 131072         hw.l1dcachesize: 65536
             | hw.l2cachesize: 4194304
        
               | my123 wrote:
               | sysctl on m1 contains the cache sizes for the little
               | cores (since those are CPUs 0-3)
               | 
               | big cores (CPU4-7) have 192KB L1I and 128KB L1D.
        
               | Nokinside wrote:
               | M1 cache lines are double of what is commonly used by
               | Intel, AMD and other ARM microarchtectures use. That's
               | significant difference.
        
               | [deleted]
        
               | JonathonW wrote:
               | Compared to the i9-9880H in my 16" MacBook Pro:
               | hw.cachelinesize: 64         hw.l1icachesize: 32768
               | hw.l1dcachesize: 32768         hw.l2cachesize: 262144
               | hw.l3cachesize: 16777216
               | 
               | The M1 doubles the line size, doubles the L1 data cache
               | (i.e. same number of lines), quadruples the L1
               | instruction cache (i.e. double the lines), and has a 16x
               | larger L2 cache, but no L3 cache.
        
       | waterside81 wrote:
       | For people who know more about this stuff than me: are these
       | sorts optimizations only possible because Apple controls the
       | whole stack and can make the hardware & OS/software perfectly
       | match up with one another or is this something that Intel can do
       | but doesn't for some reasons (tradeoffs)?
        
         | viktorcode wrote:
         | There's at least two M1 optimisations targeting Apple's
         | software stack:
         | 
         | 1. Fast uncontended atomics. Speeds up reference counting which
         | is used heavily by Objective-C code base (and Swift). Increase
         | is massive comparing to Intel.
         | 
         | 2. Guaranteed instruction ordering mode. Allows for faster Arm
         | code to be produced by Rosetta when emulating x86. Without it
         | emulation overhead would be much bigger (similar to what
         | Microsoft is experiencing).
        
         | [deleted]
        
         | [deleted]
        
         | AnthonyMouse wrote:
         | > are these sorts optimizations only possible because Apple
         | controls the whole stack and can make the hardware &
         | OS/software perfectly match up with one another or is this
         | something that Intel can do but doesn't for some reasons
         | (tradeoffs)?
         | 
         | Interestingly it's the other way around. Apple is using TSMC's
         | 5nm process (they don't have their own fabs), which is better
         | than Intel's in-house fabs, so it's _Intel 's_ vertical
         | integration which is _hurting_ them compared to the non-
         | vertically integrated Apple.
         | 
         | Also, the answer to "is this only possible because of vertical
         | integration" is always _no_. Intel and Microsoft regularly
         | coordinate to make hardware and software work together. Intel
         | is one of the largest contributors to the Linux kernel, even
         | though they don 't "own" it. Two companies coordinating with
         | one another can do anything they could do as an individual
         | company.
         | 
         | Sometimes the efficiency of this is lower because there are
         | communication barriers and isn't a single chain of command. But
         | sometimes it's higher because you don't have internal politics
         | screwing everything up when the designers would be happy with
         | outsourcing to TSMC because they have a competitive advantage,
         | but the common CEO knows that would enrich a competitor and
         | trash their internal investment in their own fabs, and forces
         | the decision that leads to less competitive products.
        
           | cma wrote:
           | Not quite vertical integration, but TSMC's 5nm fabs are
           | Apple's fabs. (exclusively for a period of time)
           | 
           | During the iPod era, Toshiba's 1.8in HD production was
           | exclusively Apple's only for music players, but Apple gets
           | all the 5nm output from TSMC for a period of time.
        
           | hinkley wrote:
           | Integration is a petri dish. It can speed up both growth and
           | decay, and it is indifferent to which one wins.
        
         | alblue wrote:
         | Sort of; Intel and AMD are stuck with the variable width
         | instruction isa that exists due to historical evolution. To do
         | something different you need a new isa.
         | 
         | Intel tried this with Itanium a while back and failed because
         | it is difficult to get software developers to target a new isa
         | and provide compilers and compiled code for everything unless
         | you use a translation layer.
         | 
         | Apple is one step ahead here because their compilers already
         | supported ARM isa (because iPhones use them) and had both the
         | OS and apps ready to go from day one of availability.
         | 
         | They also had translation technology that would allow mutating
         | x86_64 code to ARM64 code so that old apps would (on the whole)
         | run acceptably fast on the new chip.
         | 
         | To do the latter properly, Apple had to create a special mode
         | to run the arm chip with total store order for memory writes,
         | which is not standard on arm. (It would be a lot slower if they
         | didn't have that when running Rosetta translated code.)
         | 
         | So both the OS being available, and the OS influencing the ARM
         | tweaks (eg TSO) could they pull it off.
         | 
         | They also have the position that they build hardware that uses
         | those chips so can mass produce - and in fact, replace -
         | existing hardware.
         | 
         | Each of these things could be done in isolation by
         | Intel/Windows/Apps but it would be difficult to do all three.
         | 
         | Even getting JavaScript maths in a special instruction was
         | difficult enough on Intel, and that was something of benefit to
         | any browser.
         | 
         | My guess is you'll see Intel and AMD offering Arm chips in the
         | near future, as both AWS (graviton) and Apple have shown the
         | way to a new ARM future.
        
           | pjmlp wrote:
           | Intel only failed, because AMD exists and had a license to
           | produce x86 based CPUs.
           | 
           | Without AMD everyone would eventually be dragged into
           | adopting Itanium.
        
         | wmf wrote:
         | No, there's no cross-stack optimization here. The M1 gives very
         | high performance for all code.
        
           | qeternity wrote:
           | I think this gets lost in the fray between the "omg this is
           | magic" and then the Apple haters. The M1 is a very good chip.
           | Apple has hired an amazing team and resourced them well. But
           | from a pure hardware perspective, the M1 is quite
           | evolutionary. However the whole Apple Silicon experience is
           | revolutionary and magical due to the tight software pairing.
           | 
           | Both teams deserve huge praise for the tight coordination and
           | unreal execution.
        
             | acdha wrote:
             | I think this is part of the reason where there are so many
             | people trying to find reasons to downplay it: humans love
             | the idea of "one weird trick" which makes a huge difference
             | and we sometimes find those in tech but rarely for mature
             | fields like CPU design. For many people, this is
             | unsatisfying like asking an athlete their secret, and
             | getting a response like "eat well, train a lot, don't give
             | up" with nary a shortcut in sight.
        
       | fctorial wrote:
       | I can't find any info about the memory bus of apple m1. Is it 8
       | channels 16 bit each? That's drastically different from AMDs 2
       | channels 64 bit each.
       | 
       | It looks like apple m1 is much less eager when caching memory
       | rows. Maybe because it doesn't have l3 cache.
       | 
       | Edit: This test utilizes the 8x16bit memory bus of apple m1
       | fully. It's mostly just fetching random locations from memory,
       | which can all be parallelized by the cpu pipeline. It explains
       | why the results are exactly 4x slower on my ryzen 3 with 2 memory
       | channels.
       | 
       | So the summary is that m1 is optimized for dynamic languages that
       | tend to do a DDOS attack on RAM with a lot of random memory
       | access, but it might take a performance hit with compiled
       | languages and traditional HPC techniques that tend to process
       | data in sequence like ECS.
        
         | Dylan16807 wrote:
         | That would make sense for LPDDR4 but it apparently claims to
         | have a 128 byte cache line size and I'm not sure how to square
         | that with 16 bit channel width.
        
         | skohan wrote:
         | That's an interesting observation - as someone who's built a
         | few ECS implementations, one of the things I've always taken
         | for granted is that things like cache line size are more or
         | less set in stone, given the ubiquity of x86, so it's
         | interesting to consider that the rise of ARM might create
         | additional complexities there.
         | 
         | I'm a bit of two minds about this: on the one hand, for a long
         | time I've wanted a language for writing allocators which is
         | more explicit about memory, and offers good abstractions for
         | low-level memory operations (maybe Zig is going in this
         | direction). In some sense, it feels like the move towards
         | programmers thinking less about memory management has been a
         | bit of a dead-end, and what we really want is better tools for
         | memory management. Fragmentation in terms of how processors
         | handle memory goes against this goal in some ways.
         | 
         | On the other hand, it's a bit of a "holy grail" to imagine a
         | hardware stack which obviates the need for memory optimization,
         | and really does treat loading from and storing to memory
         | anywhere on the heap as the same. But I imagine that the
         | interesting things which the M1 is doing with memory are
         | helping a lot with the worst case performance, and maybe even
         | average case performance, but they're probably not doing much
         | for the best case.
        
       | gokulkrishh09 wrote:
       | Interesting article.
        
       | djacobs7 wrote:
       | Is the article saying that the M1 is slower than we would have
       | expected in this case?
       | 
       | My understanding, based on the article, is that a normal
       | processor, we would have expected arr[idx] + arr[idx+1] and
       | arr[idx] to take the same amount of time.
       | 
       | But the M1 is so parallelized that it goes to grab both arr[idx]
       | and arr[idx+1] separately. So we have to wait for both of those
       | two return. Meanwhile, on a less parallelized processor, we would
       | have done arr[idx] first and waited for it to return, and the
       | processor would realize that it already had arr[idx+1] without
       | having to do the second fetch.
       | 
       | Am I understanding this right?
        
         | phkahler wrote:
         | >> My understanding, based on the article, is that a normal
         | processor, we would have expected arr[idx] + arr[idx+1] and
         | arr[idx] to take the same amount of time.
         | 
         | That depends. If the two accesses are on the same cache line,
         | then yes. But since idx is random that will not happen
         | sometimes. He never says how big array[] is in elements or what
         | size each element is.
         | 
         | I thought DRAM also had the ability to stream out consecutive
         | addresses. If so then it looks like Apple could be missing out
         | here.
         | 
         | Then again, if his array fits in cache he's just measuring
         | instruction counts. His random indexes need to cover that whole
         | range too. There's not enough info to figure out what's going
         | on.
        
           | SekstiNi wrote:
           | > There's not enough info to figure out what's going on.
           | 
           | If you only look at the article this is true. However, the
           | source code is freely available:
           | https://github.com/lemire/Code-used-on-Daniel-Lemire-s-
           | blog/...
        
             | mrob wrote:
             | I tried it on my old (2009) 2.5GHz Phenom II X4 905e (GCC
             | 10.2.1 -O3, 64 bit) and got results almost perfectly
             | matching the conventional wisdom:                 two  :
             | 97.4 ns       two+  : 97.9 ns       three: 145.8 ns
        
             | egnehots wrote:
             | TLDR: he is using a random index with a big enough array
        
             | [deleted]
        
             | africanboy wrote:
             | I ran the benchmark on my system
             | 
             | It's a 6 years old system, fastest times are in the 25ns
             | range
             | 
             | - 2-wise+ is 5% slower than 2-wise
             | 
             | - 3-wise is 46% slower than 2-wise
             | 
             | - 3-wise is 39% slower than 2-wise+
             | 
             | on the M1
             | 
             | - 2-wise+ is 40% slower than 2-wise
             | 
             | - 3-wise is 46% slower than 2-wise
             | 
             | - 3-wise is 4% slower than 2-wise+
        
               | SekstiNi wrote:
               | Interesting, I ran it on my laptop (i7-7700HQ) with the
               | following results:
               | 
               | - 2-wise+ is 19% slower than 2-wise
               | 
               | - 3-wise is 48% slower than 2-wise
               | 
               | - 3-wise is 25% slower than 2-wise+
               | 
               | However, as mentioned in the post the numbers can vary a
               | lot, and I noticed a maximum run-to-run difference of
               | 23ms on two-wise.
        
               | nottorp wrote:
               | Ouch. This is on my $2500 i5 mbpro from 2018.
               | 
               | $ ./two_or_three
               | 
               | N = 1000000000, 953.7 MB
               | 
               | starting experiments.
               | 
               | two : 53.3 ns
               | 
               | two+ : 60.1 ns
               | 
               | three: 78.6 ns
               | 
               | bogus 1375316400
               | 
               | ------
               | 
               | 2+ 12% slower than 2
               | 
               | 3-wise 47% slower than 2
               | 
               | 3-wise 30% slower than 2+
               | 
               | -------
               | 
               | Ratios aside, that's an interesting speed leap when the
               | article gets 9 ms for 2-wise. Mind, the laptop had lots
               | of applications running, i didn't clear it up to do a
               | proper benchmark, but still.
        
             | phkahler wrote:
             | He's only got 3 million random[] numbers. Weather that's
             | enough depends on the cache size. It also bothers me to
             | read code like this where functions take parameters (like
             | N) and never use them.
        
           | eloff wrote:
           | He mentioned it's a 1GB array, and the source code is
           | available.
        
             | phkahler wrote:
             | That array is indexed by an array of random numbers and
             | there are only 3M of them. That should be enough assuming
             | even 4 bytes per index it will just fit in the 12MB cache,
             | but then there are accesses to the big array as well.
        
         | jayd16 wrote:
         | Its a little confusing because they're conflating the idea that
         | you almost certainly read at least the entire word (and not a
         | single byte) at a time with the other idea that you could fetch
         | multiple words concurrently.
        
           | duskwuff wrote:
           | Any cached memory access is going to read in the entire cache
           | line -- 64 bytes on x86, apparently 128 on M1. This is true
           | across most architectures which use caches; it isn't specific
           | to M1 or ARM.
        
             | kzrdude wrote:
             | (As I learned from recent Rust concurrency changes) on
             | newer Intel, it usually fetches two cache lines so
             | effectively 128 bytes while AMD usually 64 bytes. That's
             | the sizes they use for "cache line padded" values (I.e
             | making sure to separate two atomics by the fetch size to
             | avoid threads invalidating the cache back and forth too
             | much).
        
               | alblue wrote:
               | To be clear here, it fetches two cache lines but it
               | doesn't put the second in exclusive state until it's
               | written to; the unit of granularity is still 64b. In a
               | scanning read mode you will see the benefit but you won't
               | see the contention on writes. (The contention will come
               | from subsequent reads on that cache line though)
        
             | jayd16 wrote:
             | Yes almost certainly more than the word will be read but it
             | varies by architecture. I would think almost by definition
             | no less than a word can be read so I went with that in my
             | explanation.
        
       | syntaxing wrote:
       | I'm super curious if it's true that my 8GB M1 will die quickly
       | because of the aggressive swaps. I guess time will tell.
        
         | acdha wrote:
         | FWIW, I have a 2010 MBA which was _heavily_ used for years as a
         | primary development system. The SSD only started to show signs
         | of degraded performance last year and that wasn't massive. I
         | would be quite surprised if the technology has become worse.
        
       | lincpa wrote:
       | Apple M1 chip adopts Warehouse/Workshop Model
       | 
       | - Warehouse: unified memory
       | 
       | - Workshop: CPU, GPU and other cores
       | 
       | - Product( material): information,data
       | 
       | there's also a new unified memory architecture that lets the CPU,
       | GPU, and other cores exchange information between one another,
       | and with unified memory, the CPU and GPU can access memory
       | simultaneously rather than copying data between one area and
       | another. Accessing the same pool of memory without the need for
       | copying speeds up information exchange for faster overall
       | performance.
       | 
       | reference:
       | 
       | 1. Developer Delves Into Reasons Why Apple's M1 Chip is So Fast.
       | https://www.macrumors.com/2020/11/30/m1-chip-speed-explanati...
       | 
       | 2. The Grand Unified Programming Theory: The Pure Function
       | Pipeline Data Flow with Warehouse/Workshop Model
       | https://github.com/linpengcheng/PurefunctionPipelineDataflow
        
       | [deleted]
        
       | jeffbee wrote:
       | Great practical information. Nice to see people who know what
       | they are talking about putting data out there. I hope eventually
       | these persistent HN memes about M1 memory will die: that it's
       | "on-die" (it's not), that it's the only CPU using LPDDR4X-4267
       | (it's not), or that it's faster because the memory is 2mm closer
       | to the CPU (not that either).
       | 
       | It's faster because it has more microarchitectural resources. It
       | can load and store more, and it can do with a single core what an
       | Intel part needs all cores to accomplish.
        
         | titzer wrote:
         | > it can do with a single core what an Intel part needs all
         | cores to accomplish.
         | 
         | Care to explain what you mean specifically by this?
        
           | saagarjha wrote:
           | The M1 has extremely high single-core performance.
        
             | temac wrote:
             | It is not 4 times faster than an Intel core, though...
        
               | saagarjha wrote:
               | It is in memory performance, which is what I assumed was
               | being measured here.
        
               | kllrnohj wrote:
               | How are you defining memory performance and where are
               | your supporting comparisons? This article only discusses
               | the M1's behavior, and makes no comparisons to any other
               | CPU.
        
               | FabHK wrote:
               | FWIW, I ran it on a MacBook Pro (13-inch, 2019, Four
               | Thunderbolt 3 ports), 2.4 GHz Quad-Core Intel Core i5, 8
               | GB 2133 MHz LPDDR3:                 two  : 49.6 ns  (x
               | 5.5)       two+ : 64.8 ns  (x 5.2)       three: 72.8 ns
               | (x 5.6)
               | 
               | EDIT to add: above was just `cc`. Below is with `cc -O3
               | -Wall`, as in Lemire's article:                 two  :
               | 62.8 ns  (x 7.1)       two+ : 69.2 ns  (x 5.5)
               | three: 95.3 ns  (x 7.3)
        
               | namibj wrote:
               | You _need_ to use -mnative because it otherwise retains
               | backwards compatibility to older x86.
        
               | FabHK wrote:
               | (base) Coding % cc -mnative two-three.c       clang:
               | error: unknown argument: '-mnative'            (base)
               | Coding % cc -v       Apple clang version 12.0.0
               | (clang-1200.0.32.28)       Target: x86_64-apple-
               | darwin20.2.0       Thread model: posix
        
               | astrange wrote:
               | It's spelled "-march=native" in gcc and "-arch x86_64h"
               | in clang.
               | 
               | It doesn't make much difference though, autovectorization
               | doesn't work very well and there is not a lot of special
               | optimization for newer x86 CPUs.
        
               | [deleted]
        
               | africanboy wrote:
               | there must be something wrong there, on my late 2014
               | laptop that mounts                   Type: DDR4
               | Speed: 2133 MT/s
               | 
               | I get                   two  : 27.1 ns (3x)         two+
               | : 28.6 ns (2.2x)         three: 39.7 ns (3x)
               | 
               | which is not much, considering this is an almost 6 years
               | old system with 2x slower memor
        
               | FabHK wrote:
               | Dunno, I didn't reboot and didn't close all other
               | programs (browser, editor, mail, calendar, notes,
               | editor)... Top shows
               | 
               | Load Avg: 2.36, 2.01, 1.97 CPU usage: 2.10% user, 3.39%
               | sys, 94.49% idle
        
             | titzer wrote:
             | Sure, and it has a very large out-of-order execution
             | engine, but it is not fundamentally different from what
             | other super scalar processors do. So I am curious what the
             | OP meant by that offhand comment.
        
               | jeffbee wrote:
               | One core of the M1 can drive the memory subsystem to the
               | rails. A single core can copy (load+store) at 60GB/s.
               | This is close to the theoretical design limit for DDR4X.
               | A single core on Tiger Lake can only hit about 34GB/s,
               | and Skylake-SP only gets about 15GB/s. So yes, it is
               | close to 4x faster.
        
               | titzer wrote:
               | Thanks for clarifying. But this isn't any fundamental
               | difference IMO. There isn't any functional limitation in
               | an Intel core that means it cannot saturate the memory
               | bandwidth from a single core, unless I am missing
               | something.
        
               | nkurz wrote:
               | One could argue that it's not "fundamental", but it's
               | definitely a functional limitation of the current Intel
               | cores. The memory bandwidth of a single core is hardware
               | limited by the number "Line Fill Buffers". Each buffer
               | keeps track of one outstanding L1 cacheline miss, thus
               | the number of LFB's limits the memory level parallelism
               | (MLP). "Little's Law" gives the relationship between the
               | latency, outstanding requests, and throughput. With 10
               | LFB's and the current latency of memory, it's physically
               | impossible for a single core to use all available memory
               | bandwidth, especially on machines with more than 2 memory
               | channels.
               | 
               | The M1 chip allows higher MLP, presumably because it has
               | more LFB's per core (or maybe they are using different
               | approach where the LFB's are not per-core?). I apologize
               | for using so many abbreviations. I searched to try to
               | find a better intro, but didn't find anything perfect. I
               | did come across this thread that (apparently) I started
               | several years ago at the point where I was trying to
               | understand what was happening:
               | https://community.intel.com/t5/Software-Tuning-
               | Performance/S....
        
               | AnbeSivam wrote:
               | Came across someone else mentioning the similar bandwidth
               | constraint w.r.to LFB per core a month back.
               | 
               | https://news.ycombinator.com/item?id=25221968
        
               | jeffbee wrote:
               | I agree, it's not fundamental. It is, in particular, not
               | that other popular myth, that it's "because ARM". It's
               | only that 1 core on an Intel chip can have N-many
               | outstanding loads and 1 core of an M1 can have M>N
               | outstanding loads.
        
         | titzer wrote:
         | Frankly, I find Lemire does oversimplified, poor-quality
         | control, back-of-the-envelope microbenchmarking all the time
         | that provides little to no insight other than establishing a
         | general trend. It's sophomoric and a poor demonstration about
         | how to well-controlled benchmarking that might yield useful,
         | repeatable, and transferrable results.
        
           | alecco wrote:
           | Can you give an example? I've seen Lemire correct his posts
           | on many occasions and the source code is published. I don't
           | know many blogs doing anything remotely like that.
        
             | titzer wrote:
             | Sure. He often benchmark some small C++ code on his
             | "laptop" CPU (which one exactly? microarchs matter!) and
             | then committing classic microbenchmarking pitfalls such as:
             | 
             | - benchmarking something small enough to inspect machine
             | code, but not inspecting machine code
             | 
             | - not plotting distribution, average, variance etc
             | 
             | - no attention paid to CPU frequency governor settings
             | 
             | - measuring too short a run
             | 
             | - measuring too small a dataset that it fits entirely in L1
        
         | foldr wrote:
         | >or that it's faster because the memory is 2mm closer to the
         | CPU (not that either)
         | 
         | Not to disagree with your overall point, but 2mm is a long way
         | when dealing with high frequency signals. You can't just
         | eyeball this and infer that it makes no difference to
         | performance or power consumption.
        
           | jeffbee wrote:
           | If it works, it works. There will be no observable
           | performance difference for DDR4 SDRAM implementations with
           | the same timing parameters, regardless of the trace length.
           | There are systems out there with 15cm of traces between the
           | memory controller pins and the DRAM chips. The only thing you
           | can say against them is they might consume more power driving
           | that trace. But you wouldn't say they are meaningfully
           | slower.
        
             | foldr wrote:
             | You can't just eyeball the PCB layout for a GHz frequency
             | circuit and say "yeah that would definitely work just the
             | same if you moved that component 2mm in this direction".
             | It's certainly possible to use longer trace lengths, but
             | that may come with tradeoffs.
             | 
             | >The only thing you can say against them is they might
             | consume more power driving that trace
             | 
             | Power consumption is really important in a laptop, and
             | Apple clearly care deeply about minimising it.
             | 
             | For all we know for sure, moving the memory closer to the
             | CPU may have been part of what's enabled Apple to run
             | higher frequency memory with acceptable (to them) power
             | draw.
        
         | sliken wrote:
         | The most impressive thing I've seen is that when accessed in a
         | TLB friendly fashion that the latency is around 30ns.
         | 
         | Anandtech has a graph showing this, specifically the R per RV
         | prange graph. I've verified this personally with a small
         | microbenchmark I wrote. I've not seen anything else close to
         | this memory latency.
        
           | reasonabl_human wrote:
           | Mind sharing the micro benchmark you wrote? I'm curious to
           | know how that would work
        
             | sliken wrote:
             | https://github.com/spikebike/pstream
             | 
             | It's designed to graph latency/bandwidth for 1 to N
             | threads. My 1 thread numbers match Anandtech's. Use -p 0
             | for full random, which thrashes the TLB or -p 1 to be cache
             | friendly (visit each cacheline once, but within a sliding
             | window of 1 page).
             | 
             | To see the apple results (if you have gnuplot installed):
             | ./lview results/apple-m1
        
           | tandr wrote:
           | Sorry, what would AMD's or Intel's "latest and greatest"
           | numbers for the same be?
        
             | sliken wrote:
             | Here's the M1: https://www.anandtech.com/show/16252/mac-
             | mini-apple-m1-teste...
             | 
             | Scroll down to the latency vs size map and look at the R
             | per RV prange. That gets you 30ns or so.
             | 
             | Similar for AMD's latest/greatest the Ryzen 9 5950X:
             | https://www.anandtech.com/show/16214/amd-zen-3-ryzen-deep-
             | di...
             | 
             | The same R per RV prange is in the 60ns range.
        
               | tandr wrote:
               | Thank you very much. So we are talking about doubling (or
               | halving, depending what side you are looking from) the
               | access times.
        
               | epistasis wrote:
               | Could this be coming from the page size being 4x as large
               | for Apple Silicon versus x86? I don't fully understand
               | the benchmark, but it appears to be accessing a variety
               | of pages from the same first level TLB lookup?
               | 
               | It's been a long time since I dealt with this stuff
               | (wanted to get 1GB huge pages in Linux for some huge huge
               | hash tables), so maybe I'm misunderstanding.
        
               | sliken wrote:
               | Cachelines, page sizes, and size of the TLB all play a
               | role. But with tinkering you can see those effects
               | yourself and I played with 1, 2, 4, 8, 16, and 32 "pages"
               | which I assumed were 4KB each and didn't see much
               | difference. Measured latencies do increase slowly, but
               | you expect that as the TLB becomes progressively more of
               | a bottleneck.
               | 
               | If you use a 1GB array and see full random with much
               | higher latency than a sliding window then you can be
               | pretty sure that the page size is much less than 1GB.
               | 
               | Getting the cacheline off by a factor of 2 does make a
               | small difference since you get occasional cache hits
               | instead of zero, but as long as the array tested is
               | several times larger than cache the impact is small.
               | 
               | But all in all the M1 has excellent memory bandwidth,
               | excellent latency, and shows significantly better
               | throughput on random workloads as you use more cores.
               | Normal PC desktops have 2 memory channels (even the
               | higher end i7/i9/ryzen7/ryzen9), only the $$$$
               | workstation chips like threadripper and some of the $$$$
               | Intel's have more. The little ole M1 in a mac mini,
               | starting at $700 has at least 8 memory channels. So
               | basically the M1 delivers on all fronts, larger and lower
               | latency caches, wide issue, large reorder buffers,
               | excellent IPC, and excellent power efficiency.
        
         | ed25519FUUU wrote:
         | In other words, it's better architecture. If anything this
         | makes it seem more impressive to me.
        
           | amelius wrote:
           | No, it's the same architecture but with different parameters.
           | 
           | It's like the difference between the situation where every
           | car uses 4 cylinders, and then Apple comes along and makes a
           | car with 5 cylinders.
        
             | kllrnohj wrote:
             | Your analogy was so close! It's Apple comes along and makes
             | an 8 cylinder engine. Since, you know, the other CPUs are
             | 4-wide decode and Apple's M1 is 8-wide decode :)
        
               | ben-schaaf wrote:
               | Zen 3 is effectively 8 wide with their micro-op cache.
               | Intel similarly has been 6 wide for ages.
        
         | PragmaticPulp wrote:
         | I don't understand this competition to attribute the M1's speed
         | to _one_ specific change, while downplaying all of the others.
         | 
         | M1 is fast because they optimized everything across the board.
         | The speed is the cumulative result of many optimizations, from
         | the on-die memory to the memory clock speed to the
         | architecture.
        
           | adam_arthur wrote:
           | It's fast because they optimized everything across the board,
           | and also paid for exclusive access to TSMC 5nm process.
        
         | RachelF wrote:
         | >that it's "on-die" (it's not)
         | 
         | It appears to be mounted on the same chip package.
         | 
         | Why did Apple do this if not for speed?
        
           | wmf wrote:
           | On-package memory is not faster. I suspect it is more power
           | efficient though.
        
             | rkangel wrote:
             | It makes it easier to get to a particular clock speed. The
             | geometry, interconnect lengths etc are all tightly
             | controlled, the noise is less because you're not on the
             | main PCB and you have interconnect options that aren't
             | whatever your PCB process is (e.g. commonly gold wires).
        
         | pdpi wrote:
         | This seems to be a recurring theme with the M1, and one that,
         | in a sense, actually baffles me even more than the alternative.
         | There is no "magic" at play here, it's just lots and lots of
         | raw muscle. They just seem to have a freakishly successful
         | strategy for choosing what aspects of the processor to throw
         | that muscle at.
         | 
         | Why is that strategy simultaneously remarkably efficient and
         | remarkably high-performance? What enabled/led them to make
         | those choices where others haven't?
        
           | mhh__ wrote:
           | I think it's worth saying that because AMD have only just
           | really hit their stride, Intel were under almost zero
           | pressure to improve which has really hurt them especially
           | with the process.
           | 
           | X86 is definitely a coefficient overhead, but if Intel put
           | their designs on 5nm they'd look pretty good too - Jim Keller
           | (when he was still there) hinted their offerings for a year
           | or so in the future are significantly bigger to the point of
           | him judging it to be worth mentioning so I wouldn't write
           | them off.
        
             | megap1x3ls wrote:
             | Do you know what 5nm is?
        
           | adam_arthur wrote:
           | Certainly Apple's processors are far ahead, but they're a
           | full process generation (5nm) ahead of their competitors.
           | They paid their way to that exclusive right through TSMC.
           | 
           | I'm sure they'll still come out ahead in benchmarks, but the
           | numbers will be much closer once AMD moves to 5nm. You
           | absolutely cannot fairly compare chips from different fab
           | generations.
           | 
           | I don't see many comments hammering this point home enough...
           | it's not like the performance gap is through engineering
           | efforts that are leagues ahead. Certainly some can be
           | attributed to that, and Apple has the resources to poach any
           | talent necessary.
        
             | GeekyBear wrote:
             | A node shrink gives you a choice of cutting power,
             | improving performance, or some mix of the two.
             | 
             | Apple appears to have taken the power reduction when they
             | moved to TSMC 5nm.
             | 
             | >The one explanation and theory I have is that Apple might
             | have finally pulled back on their excessive peak power draw
             | at the maximum performance states of the CPUs and GPUs, and
             | thus peak performance wouldn't have seen such a large jump
             | this generation, but favour more sustainable thermal
             | figures.
             | 
             | Apple's A12 and A13 chips were large performance upgrades
             | both on the side of the CPU and GPU, however one criticism
             | I had made of the company's designs is that they both
             | increased the power draw beyond what was usually
             | sustainable in a mobile thermal envelope. This meant that
             | while the designs had amazing peak performance figures, the
             | chips were unable to sustain them for prolonged periods
             | beyond 2-3 minutes. Keeping that in mind, the devices
             | throttled to performance levels that were still ahead of
             | the competition, leaving Apple in a leadership position in
             | terms of efficiency.
             | 
             | https://www.anandtech.com/show/16088/apple-
             | announces-5nm-a14...
        
               | justsid wrote:
               | That seems like a pretty good trade off for mobile
               | devices. They usually don't have sustained performance
               | needs, you aren't going to render a movie or do other
               | long running computational tasks. But mobile has a lot of
               | bursty power demands: Launch dozens of app many times a
               | day for example. You want your short-ish interactions
               | with your phone to be snappy.
        
             | wmf wrote:
             | From a customer's perspective it's not my problem. Everyone
             | had the opportunity to bid on that fab capacity and they
             | decided not to.
        
               | adam_arthur wrote:
               | Yeah, totally agreed. But if you read these comments,
               | they seem to be in total amazement about the performance
               | gap and not acknowledging how much of an advantage being
               | a fab generation ahead is.
               | 
               | Customers don't care, but discussion of the merits of the
               | chip should be more nuanced about this.
               | 
               | It also implies that the gap won't exist for very long,
               | as AMD will move onto 5nm soon
        
               | tandr wrote:
               | > It also implies that the gap won't exist for very long,
               | as AMD will move onto 5nm soon
               | 
               | ... yes, if there is any capacity left. Capacity for the
               | new process is a limited resource after all.
        
               | jack_pp wrote:
               | People keep pointing this out but has Intel had such
               | significant performance improvements since sandy bridge?
               | With x86 it seems that lately you would be foolish to
               | upgrade less than once every 3-4 years because the
               | difference is just not that significant
        
               | wmf wrote:
               | Over the last decade or so Apple has gone from 10x slower
               | than Intel to parity, mostly by implementing techniques
               | that were already known. Surpassing the state of the art
               | may be harder to do consistently.
        
               | adam_arthur wrote:
               | The i7-2600K (Sandy Bridge) benchmarks at ~5000 on
               | Passmark, and the i7-10700K at about 20,000. So it seems
               | they've had quite a bit of improvement. Note this is
               | going from 32nm to 14nm.
               | 
               | Intel is in a really bad place now (in a forward-looking
               | sense), primarily due to their fab process falling behind
               | TSMC and others. You can't design your way ahead while
               | using old manufacturing technology
               | 
               | https://www.cpubenchmark.net/cpu.php?cpu=Intel+Core+i7-26
               | 00K...
        
             | valuearb wrote:
             | A node shrink is going to help AMD by 15% at best, they are
             | much farther behind than that on performance per watt.
             | 
             | AMD has done mobile CPUs that look as if they are close to
             | or even ahead of the M1 in performance, but they all use 2x
             | to 4x as much power. When higher core count versions of
             | Apple Silicon are available, they will be able to have
             | double the core counts of AMD chips at the same power
             | levels.
             | 
             | And each those cores are significantly faster than
             | individual AMD cores.
        
               | fomine3 wrote:
               | Apple M1 is fully optimized for efficiency because it's
               | delivered from mobile SoC and still aimed for portable
               | laptop, meanwhile Zen3 (or Willow Cove) arch is aimed for
               | all laptop/desktop/server category so they optimized for
               | both efficiency and max performance.
               | 
               | Even though M1 is designed for efficiency, it sometimes
               | outperforms AMD/Intel for performance. That confuses the
               | story.
        
               | GeekyBear wrote:
               | AMD's mobile chips are still on Zen 2 cores, so they are
               | behind the M1 on single core performance in both integer
               | and floating point.
               | 
               | It's their desktop chips running Zen 3 cores that trade
               | blows on single core integer math, depending on which
               | benchmark you look at.
               | 
               | https://www.anandtech.com/show/16252/mac-mini-
               | apple-m1-teste...
               | 
               | Of course, on multi-core performance you can buy a
               | desktop chip with a much higher Zen 3 core count than
               | Apple offers.
        
               | valuearb wrote:
               | The Ryzen 5950x has 135 watt TDP in actual use. That's
               | roughly 8x the M1.
               | 
               | https://www.anandtech.com/show/16214/amd-zen-3-ryzen-
               | deep-di...
               | 
               | And Ryzen chips are offered with more cores, but that's
               | an extremely temporary advantage (reminds me of the
               | friend who told me not to buy Apple stock because they
               | didn't have big screen phones).
               | 
               | When Apple fits 32 Firestorm cores in a 135 watt TDP
               | package, AMD isn't going to have an answer.
        
               | adam_arthur wrote:
               | Power consumption is not linear as it relates to
               | performance. CPUs designed for the desktop are going to
               | use excessive power by design. They'll often use many
               | times more power than mobile equivalent, but only have
               | slightly better single core performance.
               | 
               | Here's a great example, Intel Core i9 (Desktop, 125w TDP)
               | vs Intel Core i7 (Laptop, 15W). Huge power difference,
               | only ~10% difference in single core clock speeds.
               | https://cpu.userbenchmark.com/Compare/Intel-
               | Core-i9-10900K-v...
        
             | inkyoto wrote:
             | The 5nm vs 7nm vs 10nm vs whatever nm narrative is highly
             | reminiscent that of clock rate flame wars raging in the
             | tech bubble over a decade ago. Back in the early aughts,
             | when the great engineering was still a thing, alternative
             | CPU designs (DEC Alpha 21264, MIPS 10k/12k, PA RISC, POWER
             | and later UltraSPARC designs) were consistently
             | outperforming any x86 design by at least an order of
             | magnitude or more whilst being clocked at 30% to 50% less
             | of the x68 designs, especially in FP operations. The
             | alternative designs explored and utilised wider and deeper
             | pipelines, bigger L1 caches and various optimisations
             | across the entire CPU arch. Every alternative CPU design
             | had something unique to offer, and that was a great thing
             | to read about and study.
             | 
             | The commoditisation of the PC hardware has driven great CPU
             | designs into an extinction. Heck, even Oracle, that are now
             | in the business of litigation for fun and a massive profit,
             | with its prodigious cash war chest has discontinued the
             | UltraSPARC architecture due to it requiring extraordinary
             | investments on multiple fronts. PC users have long been
             | forced to be content with whatever bone the CPU
             | architecture coloniser would throw at them. There appears
             | to be a resurgence of the great engineering with M1, and,
             | hopefully, that will lead to more of the thoughtful
             | engineering in medium to long term.
             | 
             | M1 is fast due to: a solid, single vision of what a modern
             | CPU should be like, continuous investment into R&D over an
             | extended period of time, a well concerted effort of the
             | engineering, design ideas reuse across multiple product
             | lines, supply chain management, and, of course, the
             | manufacturing process. Nanometers do not make for a great
             | CPU design but rather play a supporting role. If the
             | nanometers were so important, the 2017 POWER9 design
             | manufactured at a 14 nm process with a smaller L1 cache
             | would not have been able to outperform any existing x86
             | design in 2020 in both, single core and multi core (with
             | 25% to 50% lesser number of physical cores) setups? Ryzen 3
             | has narrowed the gap, but POWER9 still takes the lead and
             | POWER10 is around the corner.
             | 
             | There is a great quote by Michael Mahon, a principal HP
             | architect, in the foreword to the PA RISC 2.0 CPU
             | architecture handbook from 1995:
             | 
             |  _The purpose of a processor architecture is to define a
             | stable interface which can efficiently couple multiple
             | generations of software investment to successive
             | generations of hardware technology. Stability and
             | efficiency are the goals, and the range of software and
             | hardware technologies expected during the architecture's
             | life determine the scope for which the goals must be
             | achieved_
             | 
             |  _..._
             | 
             |  _Efficiency also has evident value to users, but there is
             | no simple recipe for achieving it. Optimizing architectural
             | efficiency is a complex search in a multidimensional space,
             | involving disciplines ranging from device physics and
             | circuit design at the lower levels of abstraction, to
             | compiler optimizations and application structure at the
             | upper levels._
             | 
             |  _Because of the inherent complexity of the problem, the
             | design of processor architecture is an iterative, heuristic
             | process which depends upon methodical comparison of
             | alternatives (<<hill climbing>>) and upon creative flashes
             | of insight (<<peak jumping>>), guided by engineering
             | judgement and good taste._
             | 
             |  _To design an efficient processor architecture, then, one
             | needs excellent tools and measurements for accurate
             | comparisons when <<hill climbing,>> and the most creative
             | and experienced designers for superior <<peak jumping.>> At
             | HP, this need is met within a cross-functional team of
             | about twenty designers, each with depth in one or more
             | technologies, all guided by a broad vision of the system as
             | a whole_.
             | 
             | Well executed holistic approach is the reason why the entry
             | level M1 is fast. We need more of <<holistic-ism>> in
             | engineering everywhere.
        
               | [deleted]
        
           | coldtea wrote:
           | > _Why is that strategy simultaneously remarkably efficient
           | and remarkably high-performance? What enabled /led them to
           | make those choices where others haven't?_
           | 
           | The things people give them complains about:
           | 
           | (a) keeping a walled garden,
           | 
           | (b) moving fast and taking the platform to new directions all
           | at once
           | 
           | (c) controlling the whole stack
           | 
           | Which means they're not beholden to compatibility with third
           | party frameworks and big players, or with their own past, and
           | thus can rely on their APIs, third party devs etc, to cater
           | to their changes to the architecture.
           | 
           | And they're not chained to the whims of the CPU vendor (as
           | the OS vendor) or the OS vendor (as the CPU vendor) either,
           | as they serve the role of both.
           | 
           | And of course they benchmarked and profiled the hell out of
           | actual systems.
        
             | jeffbee wrote:
             | Neither A nor C makes any sense, are not supported by
             | evidence. There is no aspect of the mac or macOS that can
             | be realistically described as a "walled garden". It comes
             | with a compiler toolchain and ... well, some docs. It
             | natively runs software compiled for a foreign architecture.
             | You can do whatever you want with it. It's pretty open.
             | 
             | A "walled garden" is when there is a single source of
             | software.
        
               | matthew-wegner wrote:
               | "A" does matter a bit. Builds are uploaded to the App
               | Store include bitcode, which Apple strips on
               | distribution.
               | 
               | According to docs, enabling bitcode: "Includes bitcode
               | which allows the App Store to compile your app optimized
               | for the target devices and operating system versions, and
               | may recompile it later to take advantage of specific
               | hardware, software, or compiler changes."
               | 
               | It seems quite likely they have (and probably used) the
               | capability to recompile any app on their platform to
               | benchmark real workloads against prototype silicon
               | changes.
        
               | mamp wrote:
               | Most non-Apple apps aren't from the App Store
        
               | coldtea wrote:
               | Doesn't matter for the purposes of what the parent
               | mentioned.
               | 
               | It's enough that enough of them are.
               | 
               | Plus, even if most aren't, the "short tail" people use
               | these days probably are.
        
               | ogre_codes wrote:
               | People seem to get creative about terminology. It's not
               | remotely like what I'd consider a walled garden (xBox,
               | iOS, Playstation, etc).
        
               | paulryanrogers wrote:
               | Running apps downloaded outside the store requires
               | jumping through an increasing number of hoops or vendors
               | pay to get every build signed off by a single party.
        
               | [deleted]
        
               | coldtea wrote:
               | You don't have to pay to have your apps signed off
               | (notarized).
        
               | paulryanrogers wrote:
               | I'll have to check again because if that's the case it's
               | still a walled garden, just with free admission.
        
               | astrange wrote:
               | Just right-click > Open and you can do anything you want.
        
               | perryizgr8 wrote:
               | Which is a short wall you just jumped. To get into the
               | walled garden.
        
               | [deleted]
        
               | coldtea wrote:
               | Walled garden has many meanings, depending on context.
               | 
               | macOS promotes the App Store as the source of software
               | (even if it's not the sole), and has walls like
               | notarization requirements and the Gatekeeper to prevent
               | weeds from intruding.
               | 
               | With the App Store Apple knew that there's a pool of N
               | apps that follows its guidelines, has passed internal
               | checks for API use, and can be converted quite easily to
               | a different architecture, that it could count on.
               | 
               | Their control over the platform allowed them to enforce
               | Metal and deprecate OpenGL pronto, to add a new combined
               | iOS/macOS UI libs, to introduce Marzipan.
               | 
               | They have also added stuff like Universal Binary support,
               | and most importantly Bitcode, which abstracts away parts
               | of the underlying architecture.
               | 
               | All of those where steps towards the ARM/M1 (and future
               | developments), and all were enabled via Apple's control
               | of the hardward, software, and - sure, partial - control
               | of third party apps.
        
             | anfilt wrote:
             | I will be honest as long apple keeps this walled garden
             | shenanigans going. I am not buying any of their hardware.
        
               | anfilt wrote:
               | What's with downvotes? I know some people don't mind, but
               | it's a deal breaker to me, and it something I don't want
               | to support.
               | 
               | I don't care how good their hardware is. Moreover, good
               | luck sourcing parts if the device has trouble. Apple will
               | not sell you the parts. Even if you wanted.
               | 
               | A walled garden does not make their hardware any better.
               | If anything it makes it worse. I hope for Mac users apply
               | does not clamp down further on Macs.
        
               | coldtea wrote:
               | > _What 's with downvotes?_
               | 
               | Didn't downvote, but I think it's the same as when people
               | read a "letter to the editor" of yore, declaring that
               | some person "cancelled their subscription" because of
               | something in the magazine.
               | 
               | A natural response is "Don't let the door hit you on your
               | way out", which on HN might be expressed through a
               | downvote by some.
               | 
               | > _I don 't care how good their hardware is. Moreover,
               | good luck sourcing parts if the device has trouble. Apple
               | will not sell you the parts. Even if you wanted._
               | 
               | Well, they repair all kinds of parts, and have guarantees
               | and guarantee extension programs. But in any case, their
               | allure was never "can find parts to build my own / repair
               | damages forever" or in their stuff being cheap to own or
               | fix/replace.
               | 
               | > _A walled garden does not make their hardware any
               | better._
               | 
               | Well, it does in a few ways. Mandating how the software
               | is made, and what software is sold, when it should adapt
               | new libs to continue being sold, etc, means that they can
               | move the platform in different ways faster.
        
               | anfilt wrote:
               | > "Don't let the door hit you on your way out"
               | 
               | I never bother to enter apple's tyrannical ecosystem . So
               | there is no door to hit me on the way out.
               | 
               | > Well, they repair all kinds of parts, and have
               | guarantees and guarantee extension programs.
               | 
               | You can not get a lot surface mount chips to repair a mac
               | book or iPhone without having to look on the gray-market.
               | This is even before possible firmware issue if you manage
               | to find parts. Heck even getting full replacement boards
               | is basically impossible, unless they are from donor
               | machines that have other problems.
               | 
               | > "Well, it does in a few ways. Mandating how the
               | software is made, and what software is sold, when it
               | should adapt new libs to continue being sold, etc, means
               | that they can move the platform in different ways
               | faster."
               | 
               | I disagree, allowing people to side-load does not stop
               | apple from having policies in place for it's app stores.
               | That's honestly the only problem. It's the owners
               | hardware, they should not need apples permission to run
               | code on it. Unless the owner can sign software themselves
               | and/or run it without apple's consent this will always be
               | a problem. You can't even install gcc without
               | jailbreaking an iPhone.
               | 
               | At the very least I should be able to install an other OS
               | on the device like GNU/Linux. If apple does not want to
               | open iOS the user should at least have that option for
               | the hardware.
               | 
               | This is even before you get into how apple treats
               | developers. Have you read the entire App store
               | guidelines. Some of it is ridiculous. Some of the
               | insanity prevents Firefox from even porting their own
               | browser engine.
        
             | fartcannon wrote:
             | They could still do all this shit without the walled
             | garden. To me, it suggests they aren't willing to compete.
             | They're anti-competitive.
        
               | marrvelous wrote:
               | With the walled garden, Apple can set enforceable
               | timelines for the software ecosystem to adopt to
               | architectural changes.
               | 
               | Remember the transition to arm64? Apple forced everything
               | on the App Store to ship universal binaries.
               | 
               | Without the App Store walled garden, software isn't
               | required to keep up to date with architectural changes.
               | Instead, keeping current is only a requirement to being
               | featured on the App Store (which would just be a single
               | way to install software, not the only method).
        
               | danaris wrote:
               | Well, and on the Mac, it's _not_ the only method. The
               | walled garden here has big open gates.
               | 
               | That said, _all_ software on the Mac, post-Catalina, has
               | to be 64-bit, whether it 's distributed through the Mac
               | App Store or not, because the 32-bit system libraries are
               | no longer included at all.
        
               | astrange wrote:
               | 32-bit Windows software is actually supported through
               | WINE and works in Rosetta.
        
               | coldtea wrote:
               | This is about 32-bit OS X software/libs.
        
               | coldtea wrote:
               | > _Well, and on the Mac, it 's not the only method. The
               | walled garden here has big open gates._
               | 
               | Gates are not incompatible with walled gardens. Most
               | walled gardens have those.
               | 
               | Plus, I mentioned the walled garden as a good thing. It's
               | part of the Apple proposition (even if not alll get it),
               | and part of what it enables it to move at the speed it
               | does (whether in the right or wrong direction).
               | 
               | But one can susbstitute "walled garden" with "tight
               | control of the OS, hardware and imposed requirements on
               | most of third party software, and willingness to enforce
               | hard schedules (e.g. regarding removing 32-bit, OpenGL,
               | etc) to all (or tons) of its developers at once.
        
               | ogre_codes wrote:
               | They are little tiny 6" tall walls that you can step
               | over. Like micro-walls. Except for the bits where you
               | have no walls at all. Like if you install literally any
               | programming language, HomeBrew, or MacPorts.
               | 
               | The walls in the walled garden only exist in the heads of
               | people who never use a Mac.
        
               | coldtea wrote:
               | > _They could still do all this shit without the walled
               | garden_
               | 
               | With much slower adoption, pushback, and bike-shedding,
               | like in the Microsoft and Linux world.
               | 
               | > _To me, it suggests they aren 't willing to compete._
               | 
               | Compete with what? With themselves? They compete with
               | Windows (and to a degree Linux, though few care for
               | that), and with Android. They'd compete with Windows
               | Phone too if MS wasn't incompetent.
               | 
               | But they didn't do anything to preclude others from
               | making their OS/hardware and selling it to customers. In
               | fact, they have nowhere near a monopoly in either the
               | desktop (10% or less) or the mobile space (40% or less).
               | 
               | Whereas MS for example, had 98% of the desktop (home and
               | enterprise), and abused its power to threaten OEMs to do
               | its bidding against Linux etc.
        
               | ogre_codes wrote:
               | > They could still do all this shit without the walled
               | garden.
               | 
               | They do. MacOS isn't a walled garden.
               | 
               | > They're anti-competitive
               | 
               | Have you heard of this little company from Washington
               | called Microsoft? They have something like 85% of the PC
               | market. There is another OS called Linux. About 85-90% of
               | the internet runs on it.
               | 
               | I can understand a little where people get the idea the
               | iPhone is anti-competitive, but we're talking about MacOS
               | here.
        
               | fartcannon wrote:
               | It's the same cowardly leadership that stewards both iOS
               | and OSX.
               | 
               | Ask Amphetamine about how open they are.
        
               | coldtea wrote:
               | > _Ask Amphetamine about how open they are._
               | 
               | That's neither here nor there.
               | 
               | (a) Amphetamine could still be sold outside the Mac App
               | Store.
               | 
               | (b) An app name could be problematic even in FOSS land.
               | It's just that instead of Amphetamine being the name that
               | causes it, it will be something else. E.g. with the trend
               | of banning/changing terms like "master" (as in
               | replication primary master, not as in the owner of
               | slaves), unfortunately named apps could be thrown out
               | something or ask to be renamed to be included in a
               | distros package manager or a project.
        
               | astrange wrote:
               | There are a lot of juvenile named open source projects
               | that will definitely get in trouble or already have. GIMP
               | and LAME are examples.
        
               | ogre_codes wrote:
               | I do wish people would keep the quasi-religious aspects
               | out of things.
               | 
               | Amphetamine is a perfect example of how the Mac _isn 't_
               | a walled garden. They always had the _option_ top sell
               | outside the App Store. That is fundamentally the
               | difference between what makes a platform a walled garden.
               | They might have lost some sales because they couldn 't
               | participate in the Mac App Store, but they could still
               | sell their product. Some companies choose to avoid the
               | Mac App Store because they don't like Apple's policies.
        
           | baybal2 wrote:
           | > There is no "magic" at play here, it's just lots and lots
           | of raw muscle. They just seem to have a freakishly successful
           | strategy for choosing what aspects of the processor to throw
           | that muscle at.
           | 
           | There is no freakishly successful strategy at play there as
           | well. It's just all previous attempts at "fast ARM" chip were
           | rather half hearted "add a pipeline step there, add extra
           | register there, increase datapath width there," and not to
           | squeeze it to the limit.
        
           | smoldesu wrote:
           | The big "enabler" was their mass-purchase of 5nm lithography
           | across the board. Even still though, 4ghz*8c isn't anything
           | new, and isn't really that remarkable besides the low TDP
           | (which is incidentally dwarfed by the display, which draws up
           | to 5x more power than the CPU does). I think the big issue is
           | that Apple has painted themselves into a corner here: ARM
           | won't play nice with the larger CPUs they want to make, and
           | the pressure for them to provide a competent graphics
           | solution on custom silicon is mounting. They spent a lot of
           | time this generation marketing their "energy efficiency" and
           | battery life, but many consumers/professionals (myself
           | included) don't really care about either of these things.
        
           | barkingcat wrote:
           | The answer is that they have raw hard numbers from the
           | hundres of millions of iPads/iPhones sold each year, and can
           | use the metrics from those devices to optimize the next
           | generation of devices.
           | 
           | These improvements didn't come from nowhere. It came from
           | iterations of iOS hardware.
        
           | TYPE_FASTER wrote:
           | Apple has been iterating on their proprietary mobile ARM-
           | based processors since 2010, and has gotten really good at
           | it. I would imagine that producing billions of consumer
           | devices with these chips has helped give them a lot of
           | experience in shortened time frame.
           | 
           | I also wonder if having the hardware and software both worked
           | on in-house is an advantage. I mean, if you're developing
           | power management software for a mobile OS, and you're using a
           | 3rd-party vendor, then you read the documentation, and work
           | with the vendor if you have questions. If it's all internal,
           | you call them, and could make suggestions on future processor
           | design too based on OS usage statistics and metrics.
        
             | simonh wrote:
             | In fact there is clear evidence of this with M1. It has
             | optimised instructions to speed up retain/release of
             | NSObject subclasses, which is a frequent operation on
             | almost all Objective-C and Swift classes. They also
             | designed the M1 to support a memory management profile used
             | by x86 (and not ARM) to accelerate Rosetta translated
             | binaries. I'm sure there are more.
        
           | jandrese wrote:
           | It seems like Apple listened when people talked about how all
           | modern processors bottleneck on memory access and decided to
           | focus heavily on getting those numbers better.
           | 
           | Of course this leads to the question that if everyone in the
           | industry knew this was the issue why weren't Intel and AMD
           | pushing harder on it? They already both moved the memory
           | controller onboard so they had the opportunity to
           | aggressively optimize it like Apple has done, but instead we
           | have year after year where the memory lags behind the
           | processor in speed improvements, to the point where it is
           | ridiculous how many clock cycles a main memory access takes
           | on a modern x86 chip.
        
             | proverbialbunny wrote:
             | My guess is it had to do with limitations tied to the
             | x86_64 instruction set. It doesn't matter how much
             | modifications you do, if you don't start with a good
             | foundation, you're going to be limited to that foundation.
        
               | nkurz wrote:
               | I think the current consensus among experts is that the
               | instruction set is not the limiting factor. Modern x64
               | microprocessors have a separate front-end that handles
               | instruction decoding. These instructions are decoded to
               | internal proprietary "micro-ops". The internal buffers
               | and actual execution units see only these uops. One can
               | measure where the bottlenecks are, and it's rare to find
               | that the front-end is the bottleneck. While it's arguably
               | true that x64 is a poorly designed "foundation", it's
               | unlikely to be causing any performance difference here.
        
               | wtallis wrote:
               | > One can measure where the bottlenecks are, and it's
               | rare to find that the front-end is the bottleneck.
               | 
               | Part of this is due to the fact that x86 processor
               | designers won't include more execution units than they
               | can feed from their instruction decoders. Apple's
               | processors are much wider than x86 on both the decode and
               | execution resources, and it's pretty clear that the M1
               | would not perform as well if its decoders were as narrow
               | as current x86 cores.
        
               | nkurz wrote:
               | > it's pretty clear that the M1 would not perform as well
               | if its decoders were as narrow as current x86 cores
               | 
               | This would imply that it's able to sustain ILP greater
               | than 4 (or maybe 5 with macro-fusion). Does it actually
               | manage to do this often? If so, that's really impressive.
               | I was guessing that most of the advantage was coming from
               | the improved memory handling, and possibly a much bigger
               | reorder buffer to better take advantage of this, but I'm
               | happy to be shown otherwise.
        
               | astrange wrote:
               | There are real differences in processors caused by their
               | ISAs - it's not true that decoders mean it's all the same
               | RISC in the backend.
               | 
               | For instance, it's hard to combine instructions together,
               | which is actually an advantage for x86 (the complex
               | memory operands come for free). But it also guarantees
               | memory ordering that ARM doesn't which is a drawback.
               | 
               | I'm not sure how important this is in practice.
        
               | nkurz wrote:
               | > For instance, it's hard to combine instructions
               | together, which is actually an advantage for x86 (the
               | complex memory operands come for free).
               | 
               | True, although I just looked at the ARM assembly for
               | Daniel's example, and it's making good use of "ldpsw" to
               | load two registers from consecutive memory with a single
               | instruction. So in this particular case, it may be a
               | wash.
               | 
               | > But it also guarantees memory ordering that ARM doesn't
               | which is a drawback.
               | 
               | Yes, I wasn't considering the memory model to be part of
               | the instruction set. I agree that in general this could
               | be a big difference in performance, although I don't
               | think it comes up in Daniel's example.
               | 
               | I added a comment to Daniel's blog with my guess as to
               | what's happening to cause the observed timings in his
               | example. Feedback from anyone with better knowledge of M1
               | would be appreciated.
        
             | wmf wrote:
             | The Apple, Intel, and AMD memory controllers all look
             | pretty similar in performance to me. Memory latency is the
             | same at ~100 ns; Firestorm is clocked lower so latency is
             | lower in terms of cycles. One Firestorm core can saturate
             | the memory controller while Intel/AMD can't so that should
             | be an advantage for single-threaded scenarios. Intel/AMD
             | are behind, but I wouldn't say embarrassingly so and they
             | haven't been lazy.
        
             | sitkack wrote:
             | Because they both use new memory standards to force a
             | refresh in their CPU platforms to cause more churn and
             | revenue.
        
             | fctorial wrote:
             | > but instead we have year after year where the memory lags
             | behind the processor in speed improvements
             | 
             | That's what cache and tlb are for.
        
           | acdha wrote:
           | > What enabled/led them to make those choices where others
           | haven't?
           | 
           | Others have to some extent -- AMD is certainly not out of the
           | game -- so I'd treat this more as the question of how they've
           | been able to go more aggressively down that path. One of the
           | really obvious answers is that they control the whole stack
           | -- not just the hardware and OS but also the compilers and
           | high-level frameworks used in many demanding contexts.
           | 
           | If you're Intel or Qualcomm, you have a wider range of things
           | to support _and_ less revenue per device to support it, and
           | you are likely to have to coordinate improvements with other
           | companies who may have different priorities. Apple can
           | profile things which their users do and direct attention to
           | the right team. A company like Intel might profile something
           | and see that they can make some changes to the CPU but the
           | biggest gains would require work by a system vendor, a
           | compiler improvement, Windows/Linux kernel change, etc. --
           | they contribute a large amount of code to many open source
           | projects but even that takes time to ship and be used.
        
             | astrange wrote:
             | Intel does lots of contributions across the OS (Linux and
             | glibc) to compilers including their own (gcc, icc, ispc,
             | etc). Their problems aren't their ability, it's that Intel
             | is poorly managed and internal groups are constantly
             | fighting with each other.
             | 
             | Also, compiler support for CPUs is very overrated. Heavy
             | compiler investment was attempted with Itanium and
             | debunked; giant OoO CPUs like Intel's or M1 barely care
             | about code quality, and the compilers have very little
             | tuning for individual models.
        
               | acdha wrote:
               | > Intel does lots of contributions across the OS (Linux
               | and glibc) to compilers including their own (gcc, icc,
               | ispc, etc). Their problems aren't their ability, it's
               | that Intel is poorly managed and internal groups are
               | constantly fighting with each other.
               | 
               | I wasn't just talking about Intel but the concept of
               | separate CPU and compiler vendors in general. Intel
               | contributes a ton of open source but even if they were
               | perfectly organized it takes time for everything to
               | happen on different schedules before it's generally
               | available: get patches into something like Linux or gcc,
               | wait possibly years for Red Hat to ship a release using
               | the new version, etc. Certain users -- e.g. game or
               | scientific developers -- might jump on a new compiler or
               | feature faster, of course, but that's far from a given
               | and it means they're not going to get the across-the-
               | board excellent scores that Apple is showing.
               | 
               | > Also, compiler support for CPUs is very overrated.
               | Heavy compiler investment was attempted with Itanium and
               | debunked; giant OoO CPUs like Intel's or M1 barely care
               | about code quality, and the compilers have very little
               | tuning for individual models.
               | 
               | This isn't entirely wrong but it's definitely not
               | complete. Itanium failed because brilliant compilers
               | didn't exist and it was barely faster even with hand-
               | tuned code, especially when you adjusted for cost, but
               | that doesn't mean that it doesn't matter at all. I've
               | definitely seen significant improvements caused by CPU
               | family-specific tuning and, more importantly, when new
               | features are added (e.g. SIMD, dedicated crypto
               | instructions, etc.) a compiler or library which knows how
               | to use those can see huge improvements on specific
               | benchmarks. That was more what I had in mind since those
               | are a great example of where Apple's integration shines:
               | when they have a task like "Make H.265 video cheap on a
               | phone" or "Use ML to analyze a video stream" they can
               | profile the whole stack, decide where it makes sense to
               | add hardware acceleration, and then update their choice
               | of the compiler toolchain and higher-level libraries
               | (e.g. Accelerate.framework) and ship the entire thing at
               | the time of their choosing whereas AMD/Intel/Qualcomm and
               | maybe nVidia have to get Microsoft/Linux and maybe
               | someone like Adobe on board to get the same thing done.
               | 
               | That isn't a certain win -- Apple can't work on
               | everything at once and they certainly make mistakes --
               | but it's hard to match unless they do screw up.
        
           | SurfingInVR wrote:
           | Something I've seen no one else mentioning: Apple's low-spec
           | tier is $1000, not $70.
        
             | square_usual wrote:
             | It's $699, for a complete device, not a part of one.
        
               | perryizgr8 wrote:
               | What's the screen resolution on that $699 "complete
               | device"?
        
               | FreshFries wrote:
               | 5120 by 2880
        
               | square_usual wrote:
               | Comparable to a $70 processor, at least.
        
           | [deleted]
        
           | dv_dt wrote:
           | No fighting the sales department on where to put the market
           | segmentation bottlenecks?
        
           | gameswithgo wrote:
           | I see two main things behind it:
           | 
           | 1. they are the only ones who have 5nm chips because they
           | paid a lot to TSMC for that right 2. they gave up on
           | expandable memory, which lets them solder it right next to
           | the cpu, which likely makes it easier to ship with really
           | high clocks. and/or they just spent the money it takes to get
           | binned lpddr4 at that speed.
           | 
           | So a good cpu design, just like AMD and Intel have, but one
           | generation ahead on node size, and fast ram. Its not special
           | low latency ram or anything, just clocked higher than maybe
           | any other production machine, though enthusiasts sometimes
           | clock theirs higher on desktops!
        
             | epistasis wrote:
             | > So a good cpu design, just like AMD and Intel have
             | 
             | The design seems to be very different, in that it's far far
             | wider, and supposedly has a much better branch predictor.
             | 
             | > fast ram
             | 
             | Is that a property of the RAM clock, or a function of a
             | better memory controller? The RAM certainly doesn't appear
             | to have any better latency.
        
               | gameswithgo wrote:
               | Right, latency isn't (much) affected by a higher clock
               | rate. Getting ram to run fast requires both good ram
               | chips and good controller/motherboard.
               | 
               | and yes, obviously apples bespoke ARM cpu is quite a bit
               | different than Zen3 Ryzens x86 cpu, but I'm not sure it
               | is net-better. When Zen4 hits at 5nm I expect it will
               | perform on par or better than the M1, but we won't know
               | till it happens!
        
             | mtgx wrote:
             | In other words: money. Throwing money at the (right)
             | problems made them better than others.
             | 
             | "But doesn't Intel have a lot of money, too?"
             | 
             | Sure, but Intel has also been running around like a
             | headless chicken this past decade (pretty much literally,
             | since Otellini left) combined with them getting very
             | complacent because they had "no real competition."
        
           | gavin_gee wrote:
           | didnt they also make some interesting hires a few years ago
           | like Anand from Anandtech and some other silicon vets that
           | likely helped them design the M1 approach?
        
           | jeffbee wrote:
           | I don't have any inside-Apple perspective, but my guess is
           | having a tight feedback cycle between the profiles of their
           | own software and the abilities of their own hardware has
           | helped them greatly.
           | 
           | The reason I think so is when I was at Google is was 7 years
           | between when we told Intel what could be helpful, and when
           | they shipped hardware with the feature. Also, when AMD first
           | shipped the EPYC "Naples" it was crippled by some key uarch
           | weaknesses that anyone could have pointed out if they had
           | been able to simulate realistic large programs, instead of
           | tiny and irrelevant SPEC benchmarks. If Apple is able to
           | simulate or measure their own key workloads and get the
           | improvements in silicon in a year or two they have a gigantic
           | advantage over anyone else.
        
             | martamorena2 wrote:
             | That's bizarre. As if CPU vendors were unable to run
             | "realistic" workloads. If they truly aren't, that's because
             | they are unwilling and then they are designing for failure
             | and Apple can just eat their lunch.
        
               | kortilla wrote:
               | It's a big world out there. Workloads in data science vs
               | gaming vs hft vs packet processing vs web servers are all
               | extremely different.
               | 
               | Even if you know about them, you need an expert in each
               | to truly push the hardware to the real limits that get
               | hit in the respective industry. The small differences
               | between real implementations and simulated loads can
               | drastically alter the performance characteristics and
               | cause proc manufacturers to miss the mark.
        
               | proverbialbunny wrote:
               | As a data scientist, I feel this. Intel and AMD don't own
               | an OS or an app store, and you might be surprised how
               | hard it is to get good data. Data is the new gold. If a
               | company that can corner a piece of the market, they can
               | collect data no one else can, and from that companies are
               | often forced to partner or they can't properly provide
               | services that will keep them competitive.
        
               | pabs3 wrote:
               | Intel do own an OS, Clear Linux, but they probably lack
               | profiles of typical usage of that OS, and probably there
               | are not many users of it apart from Phoronix when they do
               | benchmarking.
               | 
               | https://clearlinux.org/
        
               | epistasis wrote:
               | This makes me think that any sort of data advantage Apple
               | may have has nothing to do with them owning an OS. Intel
               | has a massive computer network, managed by their own IT
               | team, just like any other large corporation. Intel could
               | collect whatever performance data they want from actual
               | users of actual programs just as easily as Apple could.
        
               | philistine wrote:
               | Apple is the only large company with a functional
               | organization. Could that be it ? Coupled with their
               | unparalleled ownership of a family of platforms (intel's
               | OS comparably is nonexistent).
        
               | epistasis wrote:
               | I think you're probably right, but it's funny because I
               | think Intel was well known for having really exceptional
               | organizational function in the past. They used to be
               | paranoid about everything!
        
               | [deleted]
        
               | eyelidlessness wrote:
               | Apple doesn't just control an operating system or an App
               | Store. They also control a development toolchain and the
               | two primary languages compiled for their platforms, as
               | well as most of the frameworks used in commonly used apps
               | (excepting the dreaded electron). They have a platform
               | that's been tailored to be profiled and optimized.
               | 
               | One early benchmark showed allocating and destroying an
               | NSObject performing drastically better on the M1 vs
               | recent Intel Macs. This wasn't an accident. It's probably
               | not representative of performance overall. They have
               | enough vertical integration to make their own first party
               | solutions clear optimization targets.
        
               | epistasis wrote:
               | This sort of "data," that optimizing contention free
               | locks could have big rewards, isn't something that you
               | need to control the OS or compiler/profiler/debugger
               | toolchain to understand and learn. And for that matter,
               | Intel has excellent compilers and profilers too.
               | 
               | All it takes is looking at what's going on in commonly
               | used code, deciding to optimize for X, Y, and Z, and
               | commit to it. If Intel isn't doing this already, that's
               | all the fault of current management for not making it a
               | priority.
               | 
               | The only way that Apple's vertical integration helped
               | them make that management decision is that they were able
               | to say "our customer is a typical laptop user." Intel
               | tries to cater to much larger markets, so perhaps when
               | management goes to plan a laptop chip, they are less
               | aggressive with deciding to optimize. But I have a
               | feeling that Apple's optimizations are _generally_ good
               | for nearly all code, not just for specific use cases.
        
               | jeffbee wrote:
               | There's really no explanation for why EPYC "Naples" was
               | so bad other than AMD did not internally understand the
               | performance of realistic large-scale programs. I mean
               | even if they had taken anything off the shelf, for free,
               | like MySQL, they could have determined at some point
               | before mass production that their CPU, in fact, sucked.
               | But they shipped it and prospective customers rejected
               | it.
               | 
               | Don't discount how a weak organization can make poor
               | decisions even when all necessary information seems to be
               | readily available.
        
               | [deleted]
        
             | gumby wrote:
             | > I don't have any inside-Apple perspective, but my guess
             | is having a tight feedback cycle between the profiles of
             | their own software and the abilities of their own hardware
             | has helped them greatly.
             | 
             | Also we see not the first chip but the first one that met
             | their needs (demonstrably better performance on their
             | workloads).
             | 
             | By which I mean: presumably MacOS has been running a many
             | generations of A processors, so they have had a lot of time
             | to figure out what tweaks would be good and which turn out
             | to be pessimization and overkill. It doesn't hurt that
             | there is significant internal overlap between modern macOS
             | and iOS.
        
             | dcolkitt wrote:
             | Interesting point. This would suggest pretty sizable
             | synergies from the oft-rumored Microsoft acquisition of
             | Intel.
        
               | f6v wrote:
               | > Microsoft acquisition of Intel
               | 
               | Could that possibly be approved by governments?
        
               | gigatexal wrote:
               | Nope. Not a lawyer but I doubt it at all.
        
               | ChuckNorris89 wrote:
               | Microsoft doesn't need to acquire Intel, they need to do
               | what Apple did and acquire a stellar ARM design house
               | that will build a chip with x86 translation, tailored to
               | accelerate the typical workloads on Windows machines and
               | sell those chips to the likes of Dell and Lenovo and tell
               | developers _" ARM Windows is the future, x86 Windows will
               | be sunset in 5 years and no longer supported by us, start
               | porting your apps ASAP and in the mean time, try our X86
               | emulator on our ARM silicon, it works great."_
               | 
               | Microsoft proved with the XBOX and Surface series they
               | can make good hardware if they want, now they need to
               | move to chip design.
        
               | ralfd wrote:
               | Apple has at most 10% of the computer market and is just
               | one player among many. I am skeptical Microsoft with
               | their 90% dominance would or should be allowed this much
               | power over the industry.
        
               | mechEpleb wrote:
               | 90% dominance of what is increasingly a small niche
               | market. Apple controls a large fraction of the mobile
               | device market, and everything else runs linux.
        
               | icedchai wrote:
               | The fact is outside of the tech scene, most businesses
               | and consumers runs Windows. To say this is a "small niche
               | market" is laughable. Microsoft is everywhere.
        
               | sjwright wrote:
               | The traditional personal computer market isn't nearly as
               | important as it used to be. However you slice the pie,
               | there's no way you can define the pieces to be 10% Apple
               | and 90% Microsoft with a straight face.
        
               | sitkack wrote:
               | Microsoft has a pretty good relationship with AMD from
               | the Xbox. AMD already made an Arm Opteron. Windows has
               | been multiplatform since NT 3.1 (Alpha, MIPS) and then in
               | 3.51 adding in PowerPC. You can download Windows for Arm
               | for free and run in on a Raspberry Pi.
               | 
               | Microsoft has at least one homegrown processor that it
               | has ported Windows and Linux to with the confusingly
               | named 'Edge'.
               | 
               | https://www.theregister.com/2018/06/18/microsoft_e2_edge_
               | win...
               | 
               | Microsoft doesn't even _need_ to target Arm, they could
               | easily team up with AMD or go the whole thing solo and
               | target anything from RISC-V, Arm to an in-house ISA.
        
               | mietek wrote:
               | _> Windows has been multiplatform since NT 3.1 (Alpha,
               | MIPS) and then in 3.51 adding in PowerPC._
               | 
               | What was the last version of Windows to support either of
               | these platforms?
        
               | jeffbee wrote:
               | PowerPC was the last architecture standing. It was
               | supported by NT 4.0 SP2 (technically this means Microsoft
               | supported it on paid support contracts as recently as
               | 2006).
        
               | gsnedders wrote:
               | Windows Server 2008 R2 supported Itanium.
        
               | sroussey wrote:
               | There are rumors of Ryzen CPUs with ARM ISA, but I think
               | AMD has enough on its plate.
        
               | eyelidlessness wrote:
               | From what I understand, Microsoft is excellent at
               | hardware design. They're just focused on a different
               | market (services->business rather than
               | consumers->services)
        
         | megablast wrote:
         | You'll get people guessing, since Apple itself puts out so
         | little information.
        
         | [deleted]
        
         | gameswithgo wrote:
         | What other laptop ships with LPDDR4X clocked at 4267? I agree
         | though that being closer to the cpu isn't having any
         | appreciable effect on latency, but being soldered close to the
         | cpu probably does make it easier for them to hit that high
         | clock rate.
        
           | wmf wrote:
           | Tiger Lake laptops such as the XPS 13.
        
           | jeffbee wrote:
           | As WMF mentions, Tiger Lake laptops like my Razer Book have
           | the same memory. It is not appreciably closer to the CPU in
           | the Apple design. In Intel's Tiger Lake reference designs the
           | memory is also in two chips that are mounted right next to
           | the CPU.
        
             | celrod wrote:
             | I have a Dell XPS 13 with a Tiger Lake CPU. Out of
             | curiosity, running the script:                 >
             | ./two_or_three       N = 1000000000, 953.7 MB
             | starting experiments.       two  : 30.6 ns       two+  :
             | 39.6 ns       three: 45.1 ns       bogus 1422321000
             | 
             | This is much slower than the times Lemire reported for the
             | M1. `two+` is 62% of the way between `two` and and `three`,
             | vs 88% for the M1.
             | 
             | EDIT: Adding `-march=native` didn't really change the
             | results, which makes sense given that it's a memory
             | benchmark.
        
               | gspr wrote:
               | Strange on my XPS 15 7590 (i7-9750H), I get
               | N = 1000000000, 953.7 MB        starting experiments.
               | two  : 12.8 ns        two+  : 13.7 ns        three: 19.5
               | ns        bogus 1422321000
        
               | celrod wrote:
               | Yeah, that is strange. Why was it so slow? Trying on a
               | desktop with a 7900X, I get                 N =
               | 1000000000, 953.7 MB       starting experiments.
               | two  : 17.7 ns       two+  : 19.1 ns       three: 26.4 ns
               | bogus 1422321000
               | 
               | This is again close to 50% slower than your time, but
               | nearly twice as fast. I'll try again on the laptop and
               | make sure I don't have other processes running.
        
               | gspr wrote:
               | Strange. I suppose the compiler shouldn't matter much
               | here, right? At any rate, I'm using GCC 10.2.1.
        
               | celrod wrote:
               | I'm using gcc 10.2.0. I tried clang 11 and got more or
               | less the same thing, so it doesn't seem to make much of a
               | difference.
               | 
               | Neither did messing with flags, like (I tried -fno-
               | semantic-interposition -march=native and a few others).
        
               | celrod wrote:
               | I just ran it again, and got more or less the same
               | results:                 N = 1000000000, 953.7 MB
               | starting experiments.       two  : 29.7 ns       two+  :
               | 36.5 ns       three: 43.8 ns
               | 
               | This surprises me. Normally, it does very well in most
               | benchmarks I run.
               | 
               | Looking a little closer at the script, it loads numbers
               | from "random", a vector of 3 million `Int` (this is hard
               | coded, separate from `N`). This vector is about 11.4 MiB.
               | 
               | The Tiger Lake CPU has 12 MiB of L3 cache (same as your
               | i7-9750H), so it barely fits. Meanwhile, the L1 cache is
               | 48 KiB and the L2 cache is 1.5 MiB -- huge compared to
               | most recent CPUS, and a lot of benefit in most
               | benchmarks, but at the cost of higher latency.
               | https://www.anandtech.com/show/16084/intel-tiger-lake-
               | review...
               | 
               | Skylake's L3 latency was 26-37 cycles, and in Willow
               | Cove's (Tiger Lake), it is 39-45 cycles. That difference
               | by itself isn't big enough to account for the difference
               | we're seeing, so something else must be going on.
        
               | nkurz wrote:
               | The caching of the random[] array (almost) shouldn't
               | matter, as the access is sequential.
               | 
               | I'm wondering if the difference is the the number of
               | active memory channels. How many channels does your
               | respective computers support? Do you have enough RAM
               | installed so all channels are in use? Are you able to do
               | a RAM bandwidth test by some other means to verify?
               | 
               | Another possibility is that for some reason the base
               | latency is just different between your machines. A
               | commenter added a pointer-chasing variation of Daniel's
               | test on his blog. Maybe run this to find the full latency
               | and see how the times differ?
               | 
               | Finally, there was one more commenter on the blog who
               | reported anomalously fast times on a Windows laptop. It's
               | possible there is a bug with Daniel's time measurements
               | on windows.
        
             | danaris wrote:
             | And (genuine question) how do the Tiger Lake laptops
             | compare with the M1 MacBooks thus far?
        
               | valuearb wrote:
               | The raw CPU performance of the M1 is about 15% faster
               | single core, and 50% faster multicore, while using nearly
               | half as much power.
        
               | skavi wrote:
               | AnandTech has decent benchmarks for both Tiger Lake [0]
               | and M1 [1].
               | 
               | [0]: https://www.anandtech.com/show/16084/intel-tiger-
               | lake-review...
               | 
               | [1]: https://www.anandtech.com/show/16252/mac-mini-
               | apple-m1-teste...
        
               | jeffbee wrote:
               | The outcome seems to depend greatly on the physical
               | design of the laptops. The elsewhere-mentioned Dell XPS
               | 13 has a particularly poor cooling design, which is why I
               | chose the Razer Book instead. Despite being marketed in a
               | very silly way to gamers only, it seems to have competent
               | mechanical design.
        
               | SAI_Peregrinus wrote:
               | Gamers are likely to run their systems with demanding
               | workloads, for hours, with a color-coded performance
               | counter (FPS stat). They'll notice if it throttles.
               | They're particularly demanding customers, and there's
               | quite a bit of competition for their money.
        
               | wil421 wrote:
               | How common are laptops for gamers? I always build my
               | windows boxes but I'm a casual gamer.
        
               | SAI_Peregrinus wrote:
               | I'm not sure. I know they've been getting more popular
               | with the increased power of laptops and the ability to
               | use external GPUs (via Thunderbolt 3). I'd guess desktops
               | are more common, but some people will have both.
        
             | sliken wrote:
             | Mind running a memory latency benchmark on your Razer Book?
             | Does it run linux by chance?
        
               | wmf wrote:
               | You can compare
               | https://www.anandtech.com/show/16084/intel-tiger-lake-
               | review... and https://www.anandtech.com/show/16252/mac-
               | mini-apple-m1-teste...
        
         | ksec wrote:
         | >HN memes about M1 memory will die
         | 
         | It is not only HN. It is practically the whole Internet. Go
         | around the Top 20 hardware and Apple website forum and you see
         | the same thing, also vastly amplify by a few KOL on twitter.
         | 
         | I dont remember I have ever seen anything quite like it in tech
         | circle. People were happily running around spreading
         | misinformation.
        
           | Bootvis wrote:
           | What is a KOL?
        
             | tyingq wrote:
             | "Key Opinion Leader". I think it's the new word for
             | "Influencer".
        
               | ksec wrote:
               | I am pretty sure KOL predates Influencer in modern
               | internet usage. Before that they were simply known as
               | Internet Celebrities. May be it is rarely used now. So
               | apology for not explaining the acronyms.
        
               | secondcoming wrote:
               | First I've heard of it!
        
               | walterbell wrote:
               | Who introduced the term KOL and bestows the K title?
        
           | jeffbee wrote:
           | Yeah, I know. There was some kid on Twitter who was trying to
           | tell me that it was the solder in an x86 machine (he actually
           | said "a Microsoft computer") that made them slower. Apple,
           | without the solder was much faster.
           | 
           | According to this person's bio they had an undergraduate
           | education in computer science -\\_(tsu)_/-
        
       | helloycombinat wrote:
       | I don't have a M1 chip but it'd be interesting to see what
       | happens if we make sure if the random value is never
       | n*CACHE_SIZE-1
        
       | s800 wrote:
       | What's the precision of these ns level measurements?
        
         | mhh__ wrote:
         | The answer to that is usually very context dependant, and on
         | what you're measuring. As long as you use a histogram first and
         | don't blindly calculate (say) the mean it should he obvious.
         | 
         | Two examples( that are slightly bigger than this but the same
         | principles apply):
         | 
         | If you benchmark a std::vector at insertion, you'll see a flat
         | graph with n tall spikes at ratios of it's reallocation amount
         | apart, and it scales very very well. The measurements are
         | clean.
         | 
         | If, however, you do the same for a linked list you get a
         | linearly increasing graph _but_ it 's absolutely all over the
         | place because it doesn't play nice with the memory hierarchy.
         | The std_dev of a given value of n might be a hundred times
         | worse than the vector.
        
         | CyberRabbi wrote:
         | Clock_gettime(CLOCK_REALTIME) on macos provides nanosecond-
         | level precision.
        
           | geocar wrote:
           | I seem to recall OSX didn't used to have clock_gettime, so
           | it's news to me that it even exists -- I might have been away
           | from OSX too long.
           | 
           | Is there any performance difference between that and
           | mach_absolute_time() ?
        
             | lilyball wrote:
             | It was added some years ago, and I believe
             | mach_absolute_time is actually now implemented in terms of
             | (the implementation of) clock_gettime. The documentation on
             | mach_absolute_time now even says you should use
             | clock_gettime_nsec_np(CLOCK_UPTIME_RAW) instead.
             | 
             | macOS also has clock constants for a monotonic clock that
             | increases while sleeping (unlike CLOCK_UPTIME_RAW and
             | mach_absolute_time).
        
               | saagarjha wrote:
               | Not yet, at least :)                 _mach_absolute_time:
               | 00000000000012ec        pushq   %rbp
               | 00000000000012ed        movq    %rsp, %rbp
               | 00000000000012f0        movabsq $0x7fffffe00050, %rsi
               | ## imm = 0x7FFFFFE00050       00000000000012fa
               | movl    0x18(%rsi), %r8d       00000000000012fe
               | testl   %r8d, %r8d       0000000000001301        je
               | 0x12fa       0000000000001303        lfence
               | 0000000000001306        rdtsc       0000000000001308
               | lfence       000000000000130b        shlq    $0x20, %rdx
               | 000000000000130f        orq     %rdx, %rax
               | 0000000000001312        movl    0xc(%rsi), %ecx
               | 0000000000001315        andl    $0x1f, %ecx
               | 0000000000001318        subq    (%rsi), %rax
               | 000000000000131b        shlq    %cl, %rax
               | 000000000000131e        movl    0x8(%rsi), %ecx
               | 0000000000001321        mulq    %rcx
               | 0000000000001324        shrdq   $0x20, %rdx, %rax
               | 0000000000001329        addq    0x10(%rsi), %rax
               | 000000000000132d        cmpl    0x18(%rsi), %r8d
               | 0000000000001331        jne     0x12fa
               | 0000000000001333        popq    %rbp
               | 0000000000001334        retq
        
               | Skunkleton wrote:
               | That may be the result of inlining clock_gettime, though
               | that would imply a pretty different implementation from
               | the one I am familiar with.
               | 
               | AFAIR on x86 a locked rdtsc is ~20 cycles. So to answer
               | the gp question, it has around a precision in the few
               | nanoseconds range. Accuracy is a different question, IE
               | compare numbers from the same die, but be a little more
               | suspicious across dies.
               | 
               | No clue how this is implemented on the M1, or if the M1
               | has the same modern tsc guarantees that x86 has grown
               | over the last few generations of chips.
        
               | lilyball wrote:
               | I was actually misremembering a bit.
               | 
               | Sufficiently old versions of mach_absolute_time used a
               | function called clock_get_time() on i386 (if the
               | COMM_PAGE_VERSION was not 1). This changed in macOS 10.5
               | to a tiny bit of assembly that just read from
               | _COMM_PAGE_NANOTIME on i386/x86_64/ppc (the arm
               | implementation(!!) triggers a software interrupt). The
               | i386/x86_64/ppc definitions were also copied into xnu.
               | 
               | For the next few years it kind of bounced back and forth
               | between libc and xnu, and the routine was complicated by
               | adding timebase conversion as needed. And at some point
               | arm support was added back (it vanished when it first
               | went to xnu), but this time using the commpage if
               | possible.
               | 
               | As for M1, I assume it's using the arm64 routine in xnu,
               | which can be found at https://opensource.apple.com/source
               | /xnu/xnu-7195.50.7.100.1/....
               | 
               | As for clock_gettime_nsec_np(), at least as of Big Sur,
               | it's in libc1 instead of xnu and defers to mach_continuou
               | s_time()/mach_continuous_approximate_time()/mach_absolute
               | _time()/mach_approximate_time() for the CLOCK_ __RAW[__ ]
               | clocks. And clock_gettime() for those clocks is
               | implemented in terms of clock_gettime_nsec_np().
               | 
               | 1https://opensource.apple.com/source/Libc/Libc-1439.40.11
               | /gen...
        
               | saagarjha wrote:
               | Yeah, clock_gettime is somewhat more complicated than
               | this. If anything, it might have an inlined
               | mach_absolute_time in it...
        
               | vlovich123 wrote:
               | I was part of the team that really pushed the kernel team
               | to add support for a monotonic clock that counts while
               | sleeping (this had been a persistent ask before just not
               | prioritized). We got it in for iOS 8 or 9. The dance you
               | otherwise have to do is not only complicated in userspace
               | on MacOS, it's expensive & full of footguns due to race
               | conditions (& requires changing the clock basis for your
               | entire app if I recall correctly).
        
               | lilyball wrote:
               | Do you have any insight as to why libdispatch added
               | support for this new clock internally (in the last year
               | or two), but did not expose it in any public API? In C I
               | can manually construct a dispatch_time_t that will use
               | make libdispatch use CLOCK_MONOTONIC_RAW1, if I'm willing
               | to make assumptions about the format of dispatch_time_t2
               | (despite libdispatch warning that the internal format is
               | subject to change3). And I can't even do this in Swift.
               | It would be really useful to have this functionality, so
               | I'm mystified as to why it's hidden.
               | 
               | 1Technically it uses mach_continuous_time() first if
               | available (which appears to be equivalent to
               | CLOCK_MONOTONIC_RAW), then clock_gettime(CLOCK_BOOTTIME,
               | &ts) on Linux, then clock_gettime(CLOCK_MONOTONIC, &ts),
               | then some other API for Windows.
               | 
               | 2Conveniently enough the value that is equivalent to
               | DISPATCH_TIME_NOW using the monotonic clock is just
               | INT64_MIN, at least in the current encoding.
               | 
               | 3Swift makes assumptions about the internal format of
               | dispatch_time_t so I don't know if it actually can
               | meaningfully change. Newer versions of Swift now use
               | stdlib on the system, but any app built with a
               | sufficiently old version of Swift still embeds its own
               | copy of the stdlib. Granted, additions (like the
               | monotonic clock) should be fine, as the Swift API does
               | not actually bridge from dispatch_time_t so it only ends
               | up representing times it has APIs to construct. Since it
               | doesn't have APIs to construct monotonic times, they
               | won't break it.
        
               | vlovich123 wrote:
               | I haven't worked at Apple in almost 5 years so I can't
               | speak to more recent developments unfortunately.
               | libdispatch integration may be trickier & have
               | implementation details not suitable for a public API.
               | 
               | Specifically if I recall correctly the internal
               | representation of time only allowed for 2 formats
               | (wallclock vs monotonic). I suspect adding other types of
               | clocks could pose back/forward compat challenges.
               | However, this is a wild shot in the dark & I didn't
               | really look into the internals of libdispatch. Maybe ask
               | on their github page?
               | 
               | Generally upgrading a private API to public is a lot of
               | work & goes through _a lot_ of review (having monitored
               | that mailing list  & added 1 API during my time there).
               | So if there's a private API probably some team at Apple
               | needed it to deliver a feature but the maintainers were
               | not confident the specific solution they chose
               | generalized well & either requires some work or something
               | else.
        
             | saagarjha wrote:
             | It's new in macOS Sierra. I believe mach_absolute_time is
             | slightly faster but not by much-both just read the commpage
             | these days to save on a syscall.
        
           | varjag wrote:
           | Still single digit level nanosecond precision sounds
           | marginal. 1ns = 2 clock cycles at 2GHz.
        
       | kergonath wrote:
       | It's a good introduction, but it's a bit disappointing that it
       | ends that way. I'd love to read more about what's behind the
       | figure and more technical info about how it might work.
        
         | alblue wrote:
         | This isn't specific to the M1 but I tap about cache lines in my
         | last QCon presentation (where I also suggested that a 128b
         | cache line wasn't far away):
         | 
         | https://www.infoq.com/presentations/microarchitecture-modern...
         | 
         | However the speed benefits come from a much larger L1 cache and
         | the fact that the ram is in the same chip which will reduce
         | latency that is the benefit for most of it.
         | 
         | The program (instruction) cache is also a lot bigger and has
         | the advantage that as a fixed size isa can be much wider in
         | execution than in x86 but that's unlikely to be of benefit
         | here, other than perhaps slightly in terms of queuing up
         | multiple outstanding loads.
        
           | kristianp wrote:
           | This post says that the m1 has a 128 byte cache line size. So
           | that time has arrived!
           | 
           | https://news.ycombinator.com/item?id=25660769
        
       ___________________________________________________________________
       (page generated 2021-01-07 23:02 UTC)