[HN Gopher] Python 3.15's interpreter for Windows x86-64 should ...
       ___________________________________________________________________
        
       Python 3.15's interpreter for Windows x86-64 should hopefully be
       15% faster
        
       Author : lumpa
       Score  : 286 points
       Date   : 2025-12-25 13:02 UTC (9 hours ago)
        
 (HTM) web link (fidget-spinner.github.io)
 (TXT) w3m dump (fidget-spinner.github.io)
        
       | machinationu wrote:
       | The Python interpreter core loop sounds like the perfect problem
       | for AlphaEvolve. Or it's open source equivalent OpenEvolve if
       | DeepMind doesn't want to speed up Python for the competition.
        
       | g947o wrote:
       | > This has caused many issues for compilers in the past, too many
       | to list in fact. I have a EuroPython 2025 talk about this.
       | 
       | Looks like it refers to this:
       | 
       | https://youtu.be/pUj32SF94Zw
       | 
       | (wish it's a link in the article)
        
       | Hendrikto wrote:
       | TLDR: The tail-calling interpreter is slightly faster than
       | computed goto.
       | 
       | > I used to believe the the tailcalling interpreters get their
       | speedup from better register use. While I still believe that now,
       | I suspect that is not the main reason for speedups in CPython.
       | 
       | > My main guess now is that tail calling resets compiler
       | heuristics to sane levels, so that compilers can do their jobs.
       | 
       | > Let me show an example, at the time of writing, CPython 3.15's
       | interpreter loop is around 12k lines of C code. That's 12k lines
       | in a single function for the switch-case and computed goto
       | interpreter.
       | 
       | > [...] In short, this overly large function breaks a lot of
       | compiler heuristics.
       | 
       | > One of the most beneficial optimisations is inlining. In the
       | past, we've found that compilers sometimes straight up refuse to
       | inline even the simplest of functions in that 12k loc eval loop.
        
         | kccqzy wrote:
         | I think in the protobuf example the musttail did in fact
         | benefit from better register use. All the functions are called
         | with the same arguments, so there is no need to shuffle the
         | registers. The same six register-passed arguments are reused
         | from one function to the next.
        
         | cma wrote:
         | Does MSVC support computed goto?
        
       | mishrapravin441 wrote:
       | Really nice results on MSVC. The idea that tail calls effectively
       | reset compiler heuristics and unblock inlining is pretty
       | convincing. One thing that worries me though is the reliance on
       | undocumented MSVC behavior -- if this becomes widely shipped,
       | CPython could end up depending on optimizer guarantees that
       | aren't actually stable. Curious how you're thinking about long-
       | term maintainability and the impact on debugging/profiling.
        
         | kenjin4096 wrote:
         | Thanks for reading! For now, we maintain all 3 of the
         | interpreters in CPython. We don't plan to remove the other
         | interpreters anytime soon, probably never. If MSVC breaks the
         | tail calling interpreter, we'll just go back to building and
         | distributing the switch-case interpreter. Windows binaries will
         | be slower again, but such is life :(.
         | 
         | Also the interpreter loop's dispatch is autogenerated and can
         | be selected via configure flags. So there's almost no
         | additional maintenance overhead. The main burden is the MSVC-
         | specific changes we needed to get this working (amounting to a
         | few hundred lines of code).
         | 
         | > Impact on debugging/profiling
         | 
         | I don't think there should be any, at least for Windows. Though
         | I can't say for certain.
        
           | mishrapravin441 wrote:
           | That makes sense, thanks for the detailed clarification.
           | Having the switch-case interpreter as a fallback and keeping
           | the dispatch autogenerated definitely reduces the long-term
           | risk.
        
         | pxeger1 wrote:
         | Profile of llm generated comments
        
           | mishrapravin441 wrote:
           | ust to clarify, I'm writing these comments myself. I use
           | grammar llm plugin though to clean up phrasing, but the
           | substance is mine.
        
       | develatio wrote:
       | if the author of this blog reads this: can we can an RSS, please?
        
         | kenjin4096 wrote:
         | Got it. I'll try to set one up this weekend.
        
       | redox99 wrote:
       | This seems like very low hanging fruit. How is the core loop not
       | already hyper optimized?
       | 
       | I'd have expected it to be hand rolled assembly for the major
       | ISAs, with a C backup for less common ones.
       | 
       | How much energy has been wasted worldwide because of a relatively
       | unoptimized interpreter?
        
         | kccqzy wrote:
         | Python's goal is never really to be fast. If that were its
         | goal, it would've had a JIT long ago instead of toying with
         | optimizing the interpreter. Guido prioritized code simplicity
         | over speed. A lot of speed improvements including the JIT (PEP
         | 744 - JIT Compilation) came about after he stepped down.
        
           | davidkhess wrote:
           | Should probably mention that Guido ended up on the team
           | working on a pretty credible JIT effort. Though Microsoft
           | subsequently threw a wrench in it with layoffs. Not sure the
           | status now.
        
           | IshKebab wrote:
           | If performance was a goal... hell if it was even a
           | consideration then the language would be very different.
        
           | mhh__ wrote:
           | Python is full of decisions like this / or rather full of "if
           | you just did some more work it'd be 10x better"
        
         | LtWorf wrote:
         | If you want fast just use pypy and forget about cpython.
        
         | mkoubaa wrote:
         | Software has gotten so slow we've forgotten how fast computers
         | are
        
         | pjc50 wrote:
         | This is (a) wildly over expectations for open source and (b) a
         | massive pain to maintain, and (c) not even the biggest
         | timewaster of python, which is the packaging "system".
        
           | loeg wrote:
           | > not even the biggest timewaster of python, which is the
           | packaging "system".
           | 
           | For frequent, short-running scripts: start-up time! Every
           | import has to scan a billion different directories for where
           | the module might live, even for standard modules included
           | with the interpreter.
        
             | tweakimp wrote:
             | In the near future we will use lazy imports :)
             | https://peps.python.org/pep-0810/
        
               | theLiminator wrote:
               | This can't come soon enough. Python is great for CLIs
               | until you build something complex and a simple --help
               | takes seconds. It's not something easily worked around
               | without making your code very ugly.
        
         | WD-42 wrote:
         | Probably because anyone concerned with performance wasn't
         | running workloads on Windows to begin with.
        
           | loeg wrote:
           | They weren't using Python, anyway.
        
           | pjmlp wrote:
           | Games and Proton.
           | 
           | Apparently people that care about performance do run Windows.
        
             | nilamo wrote:
             | Games are made for windows because that's where the device
             | drivers have historically been. Any other viewpoint is
             | ignoring reality.
        
               | pjmlp wrote:
               | Sure, keep believing that while loading Proton.
        
             | WD-42 wrote:
             | None of those games, or a very small amount of them, are
             | written in python. None of the ones that need to be
             | performant for sure.
        
               | pjmlp wrote:
               | Indeed, but the question was about performance in
               | general.
        
             | whatevaa wrote:
             | But not python.
        
               | pjmlp wrote:
               | Sure, but that wasn't the question.
        
         | Calavar wrote:
         | Quite to the contrary, I'd say this update is evidence of the
         | inner loop being hyperoptimized!
         | 
         | MSVC's support for musttail is hot off the press:
         | 
         | > The [[msvc::musttail]] attribute, introduced in MSVC Build
         | Tools version 14.50, is an experimental x64-only Microsoft-
         | specific attribute that enforces tail-call optimization. [1]
         | 
         | MSVC Build Tools version 14.50 was released _last month_ , and
         | it only took a few weeks for the CPython crew to turn that
         | around into a performance improvement.
         | 
         | [1] https://learn.microsoft.com/en-
         | us/cpp/cpp/attributes?view=ms...
        
       | mananaysiempre wrote:
       | The money shot (wish this were included in the blog post):
       | #   if defined(_MSC_VER) && !defined(__clang__)       #
       | define Py_MUSTTAIL [[msvc::musttail]]       #      define
       | Py_PRESERVE_NONE_CC __preserve_none       #   else       #
       | define Py_MUSTTAIL __attribute__((musttail))       #       define
       | Py_PRESERVE_NONE_CC __attribute__((preserve_none))       #
       | endif
       | 
       | https://github.com/python/cpython/pull/143068/files#diff-45b...
       | 
       | Apparently(?) this also needs to be attached to the function
       | declarator and does not work as a function specifier: `static
       | void *__preserve_none slowpath();` and not `__preserve_none
       | static void *slowpath();` (unlike GCC attribute syntax, which
       | tends to be fairly gung-ho about this sort of thing, sometimes
       | with confusing results).
       | 
       | Yay to getting undocumented MSVC features disclosed if Microsoft
       | thinks you're important enough :/
        
         | publicdebates wrote:
         | Important enough, or benefits them directly? I have no good
         | guesses how improving Python's performance would benefit them,
         | but I would guess that's the real reason.
        
           | HPsquared wrote:
           | I wonder if this is related to Python in Excel. You'll have
           | lots of people running numerical stuff written in Python,
           | running on Microsoft servers.
        
           | mkoubaa wrote:
           | A lot of commercial engineering and scientific software runs
           | on windows.
        
           | andix wrote:
           | I guess there are some Python workloads on Azure, Microsoft
           | provides a lot of data analysis and LLM tools as a service
           | (not paid by CPU minutes). Saving CPU cycles there directly
           | translates to financial savings.
        
           | acdha wrote:
           | Think about how much effort they have put into things like
           | Pylance and general python support in VAC. Clearly they think
           | they have enough users that this matters to that a first
           | class experience is worth having.
        
           | pjmlp wrote:
           | Microsoft was the one hiring Guido out of retirement, and
           | alongside Facebook finally kicking off the CPython JIT
           | efforts.
           | 
           | Python is one of the Microsoft blessed languages on their
           | devblogs.
        
             | throw1ahs wrote:
             | The project was first suggested by Mark Shannon. Van Rossum
             | inserted himself into the project. Faster CPython people
             | have been fired by Microsoft last year.
             | 
             | Generally not that much has happened in 5 years, sometimes
             | 10-15% improvements are posted that are later offset by
             | bloat.
             | 
             | I think the project started in 3.10, so 3.9 is the last
             | version to compare to. The improvements aren't that great,
             | I don't think any other language would get so much positive
             | feedback for so little.
        
               | pjmlp wrote:
               | I know what happened last year, my point was the prior
               | history that lead to that effort.
               | 
               | https://thenewstack.io/guido-van-rossums-ambitious-plans-
               | for...
               | 
               | Agree with the sentiment, Python is the only dynamic
               | language where it seems a graveyard from efforts.
               | 
               | And nope it isn't the dynamism per se, Smalltalk, Self,
               | Common Lisp are just as dynamic, with lots of
               | possibilities to reboot the world and mess up JIT
               | efforts, as any change impacts the whole image.
               | 
               | Naturally those don't have internals exposed to C where
               | anything goes, and the culture C libraries are seen as
               | the language libraries.
        
               | boulos wrote:
               | Ehh, PHP fits that bill and is clearly optimizable. All
               | sorts of things worked well for PHP, including the
               | original HipHop, HHVM, my own work, and the mainline PHP
               | runtime.
               | 
               | Python has some semantics and behaviors that are
               | particularly hostile to optimization, but as the Faster
               | Python and related efforts have suggested, the main
               | challenge is full compatibility including extensions plus
               | the historical desire for a simple implementation within
               | CPython.
               | 
               | There are limits to retrofitting truly high performance
               | to any of these languages. You want enough static,
               | optional, or gradual typing to make it fast enough in the
               | common case. That's why you also saw the V8 folks give up
               | and make Dart, the Facebook ones made Hack, etc. It's
               | telling that none of those gained truly broad adoption
               | though. Performance isn't all that matters, especially
               | once you have an established codebase and ecosystem.
        
               | gsnedders wrote:
               | > Performance isn't all that matters, especially once you
               | have an established codebase and ecosystem.
               | 
               | And this is no small part of why Java and JS have
               | frequently been pushing VM performance forward -- there's
               | enough code people very much care about continuing to
               | work on performance. (Though the two care about different
               | things mostly: Java cares much more about long-term
               | performance, and JS cares much more about short-term
               | performance.)
               | 
               | It doesn't hurt they're both languages which are
               | relatively static compared with e.g. Python, either.
        
               | kenjin4096 wrote:
               | > Generally not that much has happened in 5 years,
               | sometimes 10-15% improvements are posted that are later
               | offset by bloat.
               | 
               | Sorry but unless your workload is some C API numpy number
               | cruncher that just does matmuls on the CPU, that's
               | probably false.
               | 
               | In 3.11 alone, CPython sped up by around 25% over 3.10 on
               | pyperformance for x86-64 Ubuntu. https://docs.python.org/
               | 3/whatsnew/3.11.html#whatsnew311-fas...
               | 
               | 3.14 is 35-45% faster than CPython 3.10 for pyperformance
               | x86-64 Ubuntu https://github.com/faster-
               | cpython/benchmarking-public
               | 
               | These speedups have been verified by external projects.
               | For example, a Python MLIR compiler that I follow has
               | found a geometric mean 36% speedup moving from CPython
               | 3.10 to 3.11 (page 49 of
               | https://github.com/EdmundGoodman/masters-project-report)
               | 
               | Another academic benchmark here observed an around 1.8x
               | speedup on their benchmark suite for 3.13 vs 3.10
               | https://youtu.be/03DswsNUBdQ?t=145
               | 
               | CPython 3.11 sped up enough that PyPy in comparison looks
               | slightly slower. I don't know if anyone still remembers
               | this: but back in the CPython 3.9 days, PyPy had over 4x
               | speedup over CPython on the PyPy benchmark suite, now
               | it's 2.8 on their website https://speed.pypy.org/ for
               | 3.11.
               | 
               | Yes CPython is still slow, but it's getting faster :).
               | 
               | Disclaimer: I'm just a volunteer, not an employee of
               | Microsoft, so I don't have a perf report to answer to.
               | This is just my biased opinion.
        
             | __turbobrew__ wrote:
             | How to stay employed for life: create a programming
             | language which is pretty good, but with some fatal flaws
             | (GIL, typing, slow) and you are set for life.
        
               | johncolanduoni wrote:
               | I don't know about calling them fatal, but inculcating a
               | culture that believes the flaws are inescapable laws of
               | reality is probably key.
        
         | kenjin4096 wrote:
         | So it seems I was wrong, [[msvc::musttail]] is documented! I
         | will update the blog post to reflect that.
         | 
         | https://news.ycombinator.com/item?id=46385526
        
       | bgwalter wrote:
       | MSVC mostly generates slower code than gcc/clang, so maybe this
       | trick reduces the gap.
        
         | metaltyphoon wrote:
         | Is this backed by real evidence?
        
           | bluecalm wrote:
           | My experience is 10%-15% slower than GCC. That was 10 years
           | ago though.
        
       | jtrn wrote:
       | Im a bit out of the loop with this, but hope its not like that
       | time with python 3.14, when it was claimed a geometric mean
       | speedup of about 9-15% over the standard interpreter when built
       | with Clang 19. It turned out the results were inflated due to a
       | bug in LLVM 19 that prevented proper "tail duplication"
       | optimization in the baseline interpreter's dispatch loop. Actual
       | gains was aprox 4%.
       | 
       | Edit: Read through it and have come to the conclusion that the
       | post is 100% OK and properly framed: He explicitly says his
       | approach is to "sharing early and making a fool of myself,"
       | prioritizing transparency and rapid iteration over ironclad
       | verification upfront.
       | 
       | One could make an argument that he should have cross-compiler
       | checks, independent audits, or delayed announcements until
       | results are bulletproof across all platforms. But given that he
       | is 100% transparent with his thinking and how he works, it's all
       | good in the hood.
        
         | kenjin4096 wrote:
         | Thanks :), that was indeed my intention. I think the previous
         | 3.14 mistake was actually a good one on hindsight, because if I
         | didn't publicize our work early, I wouldn't have caught the
         | attention of Nelson. Nelson also probably wouldn't have spent
         | one month digging into the Clang 19 bug. This also meant the
         | bug wouldn't have been caught in the betas, and might've been
         | out with the actual release, which would have been way worse.
         | So this was all a happy accident on hindsight that I'm grateful
         | for as it means overall CPython still benefited!
         | 
         | Also this time, I'm pretty confident because there are two perf
         | improvements here: the dispatch logic, and the inlining. MSVC
         | can actually convert switch-case interpreters to threaded code
         | automatically if some conditions are met [1]. However, it does
         | not seem to do that for the current CPython interpreter. In
         | this case, I suspect the CPython interpreter loop is just too
         | complicated to meet those conditions. The key point also that
         | we would be relying on MSVC again to do its magic, but this
         | tail calling approach gives more control to the writers of the
         | C code. The inlining is pretty much impossible to convince MSVC
         | to do except with `__forceinline` or changing things to use
         | macros [2]. However, we don't just mark every function as
         | forceinline in CPython as it might negatively affect other
         | compilers.
         | 
         | [1]: https://github.com/faster-cpython/ideas/issues/183 [2]:
         | https://github.com/python/cpython/issues/121263
        
           | jtrn wrote:
           | I wish all self-promoting scientists and sensationalizing
           | journalists had a fraction of the honesty and dedication to
           | actual truth and proper communication of truths as you do.
           | You seem to feel that it's more important to be transparent
           | about these kinds of technical details than other people are
           | about their claims in clinical medical research. Thank you so
           | much for all you do and the way you communicate about it.
           | 
           | Also, I'm not that familiar with the whole process, but I
           | just wanted to say that I think you were too hard on yourself
           | during the last performance drama. So thank you again and
           | remember not to hold yourself to an impossible standard no
           | one else does.
        
             | kenjin4096 wrote:
             | Thank you very much for the kind words, that means a lot to
             | me!
        
             | halflings wrote:
             | +1, reading through the post, the PR updating the
             | documentation... thanks for being transparent, but also
             | don't be so hard on yourself!
             | 
             | That was a very niche error, that you promptly corrected,
             | no need to be so apologetic about it! And thanks for all
             | the hard work making Python faster!
        
         | haberman wrote:
         | I'll repeat what I said at that time: one of the benefits of
         | the new design is that it's less vulnerable to the whims of the
         | optimizer: https://news.ycombinator.com/item?id=43322451
         | 
         | If getting the optimal code is relying on getting a pile of
         | heuristics to go in your favor, you're more vulnerable to the
         | possibility that someday the heuristics will go the other way.
         | Tail duplication is what we want in case, but it's possible
         | that a future version of the compiler could decide that it's
         | not desired because of the increased code size.
         | 
         | With the new design, the Python interpreter can express the
         | desired shape of the machine code more directly, leaving it
         | less vulnerable to the whims of the optimizer.
        
           | kenjin4096 wrote:
           | Yeah, I believe that statement and it seems to hold true for
           | MSVC as well. Thanks for your work inspiring all of this btw!
        
       | acemarke wrote:
       | I've never seen this kind of benchmark graph before, and it looks
       | really neat! How was this generated? What tool was used for the
       | benchmarks?
       | 
       | (I actually spent most of Sep/Oct working on optimizing the Immer
       | JS immutable update library, and used a benchmarking tool called
       | `mitata`, so I was doing a lot of this same kind of work:
       | https://github.com/immerjs/immer/pull/1183 . Would love to add
       | some new tools to my repertoire here!)
        
         | eesmith wrote:
         | Are you referring to the violin plot?
         | https://en.wikipedia.org/wiki/Violin_plot and in Matplotlib as
         | https://matplotlib.org/stable/api/_as_gen/matplotlib.pyplot....
         | 
         | It's in essence a histogram for the distribution, with
         | smoothing, and mirrored on each side.
         | 
         | It looks nice, but is not without well-deserved opposition
         | because 1) the use of smoothing can hide the actual
         | distribution, 2) mirroring contains no extra information, while
         | taking up space, and implying the extra space contains
         | information, and 3) when shown vertically, too often causes
         | people to exclaim it looks like a vulva.
         | 
         | In an HN discussion on the topic, medstrom at
         | https://news.ycombinator.com/item?id=40766519 points to a half-
         | violin plot at
         | https://miro.medium.com/v2/1*J3Q4JKXa9WwJHtNaXRu-kQ.jpeg with
         | the histogram on the left, and the half-violin on the right,
         | which gives you a chance to see side-by-side presentation of
         | the same data.
        
           | Tarq0n wrote:
           | Histograms aren't necessarily a true depiction of the
           | distribution. Bin count or width has a large impact on what
           | details get shown.
        
       | forrestthewoods wrote:
       | Is there a Clang based build for Windows? I've been slowly moving
       | my Windows builds from MSVC to Clang. Which still uses the
       | Microsoft STL implementation.
       | 
       | So far I think using clang instead of MSVC compiler is a strict
       | win? Not a huge difference mind you. But a win nonetheless.
        
       | gozzoo wrote:
       | I have quetion - slightly off topic, but related. I was wandering
       | why is pyhton interpreter so much slower than V8 javascript
       | interpreter when both javascript and python are dynamic
       | interpreted languages.
        
         | bheadmaster wrote:
         | I can think of two possible reasons:
         | 
         | First is the Google's manpower. Google somehow succeeds in
         | writing fast software. Most Google products I use are fast in
         | contrast to the rest of the ecosystem. It's possible that
         | Google simply did a better job.
         | 
         | The second is CPython legacy. There are faster implementations
         | of Python that completely implement the API (PyPy comes to
         | mind), but there's a huge ecosystem of C extensions written
         | with CPython bindings, which make it virtually impossible to
         | break compatibility. It is possible that this legacy prevents
         | many possible optimizations. On the other hand, V8 only needs
         | to keep compatibility on code-level, which allows them to
         | practically switch out the whole inside in incremental search
         | for a faster version.
         | 
         | I might be wrong, so take what I said with a grain of salt.
        
           | canucker2016 wrote:
           | Don't forget that there was a Google attempt at making a
           | faster Python - Unladen Swallow. It got lots of PR but never
           | merged with mainline CPython (wikipedia says a dev branch was
           | released).
           | 
           | see https://en.wikipedia.org/wiki/Unladen_Swallow
        
             | pansa2 wrote:
             | Unladen Swallow got a lot of hype but was only a very small
             | project. IIRC the only people working on it were two
             | interns.
             | 
             | V8 was a _much_ higher priority - Google hired many of the
             | world's best VM engineers to develop it.
        
         | everforward wrote:
         | I know of a couple reasons offhand.
         | 
         | JavaScript is JIT'ed where CPython is not. Pypy has JIT and is
         | faster, but I think is incompatible with C extensions.
         | 
         | I think Pythons threading model also adds complexity to
         | optimizing where JavaScripts single thread is easier to
         | optimize.
         | 
         | I would also say there's generally less impetus to optimize
         | CPython. At least until WASM, JavaScript was sort of stuck with
         | the performance the interpreter had. Python had more off-ramps.
         | You could use pypy for more pure Python stuff, or offload
         | computationally heavy stuff to a C extension.
         | 
         | I think there are some language differences that make
         | JavaScript easier to optimize, but I'm not super qualified to
         | speak on that.
        
           | pansa2 wrote:
           | > _I would also say there's generally less impetus to
           | optimize CPython_
           | 
           | Nonetheless, Microsoft employed a whole "Faster CPython" team
           | for 4 years - they targeted a 5x speedup but could only
           | achieve ~1.5x. Why couldn't they make a significantly faster
           | Python implementation, especially given that PyPy exists and
           | proves it's possible?
        
             | everforward wrote:
             | Pypy has much slower C interop than CPython, which I
             | believe is part of the tradeoff. Eg data analysis pipelines
             | are probably still faster in numpy on CPython than pypy.
             | 
             | Not an expert here, but my understanding is that Python is
             | dynamic to the point that optimizing is hard. Like allowing
             | one namespace to modify another; last I used it, the
             | Stackdriver logging adapter for Python would overwrite the
             | stdlib logging library. You import stackdriver, and it
             | changes logging to send logs to stackdriver.
             | 
             | All package level names (functions and variables) are
             | effectively global, mutable variables.
             | 
             | I suspect a dramatically faster Python would involve
             | disabling some of the more unhinged mutability. Eg package
             | functions and variables cannot be mutated, only wrapped
             | into a new variable.
        
         | _ZeD_ wrote:
         | keep in mind that, apart from the money throw at js runtime
         | interpreters by google and others, there is also the fact that
         | python - as a language - is way more "dynamic" than javascript.
         | 
         | Even "simple" stuff like field access in python may refer to
         | multiple dynamically-mapped method resolution.
         | 
         | Also, the ffi-bindings of python, while offering a way to
         | extend it with libraries written in c/c++/fortran/... , limit
         | how freely the internals can be changed (see the bug-by-bug
         | compatibility work done for example by pypy, just to name an
         | example, with some constraint that limit some optimizations)
        
           | pansa2 wrote:
           | > _python - as a language - is way more "dynamic" than
           | javascript_
           | 
           | Very true, but IMO the existence of PyPy proves that this
           | doesn't necessarily prevent a fast implementation. I think
           | the reason for CPython's poor performance must be your other
           | point:
           | 
           | > _the ffi-bindings of python [...] limit how freely the
           | internals can be changed_
        
         | dragonwriter wrote:
         | > why is pyhton interpreter so much slower than V8 javascript
         | interpreter when both javascript and python are dynamic
         | interpreted languages.
         | 
         | Because JS's centrality to the web and V8's speed's centrality
         | to Google's push to avoid other platform owners controlling the
         | web via platform-default browsers meant virtually unlimited
         | resources were spent in optimizing V8 at a time when the JS
         | language itself was basically static; Python has never had the
         | same level of investment and has always spent some of its
         | smaller resources on advancing the language rather than
         | optimizing the implementation.
         | 
         | Also, because the JS legacy that needed to be supported through
         | that is _pure_ JS, whereas with CPython there is also a
         | considerable ecosystem of code that deeply integrates with
         | Python from the outside that must still be supported, and the
         | interface used by that code limits the optimizations that can
         | be applied. Faster Python interpreters exist that don't support
         | that external ecosystem, but they are less used because that
         | ecosystem is a big part of Python's value proposition.
        
         | IshKebab wrote:
         | Even though Javascript is quite dynamic, Python is _much_
         | worse. Basically everything involves a runtime look-up. It 's
         | pretty much the language you'd design if you were trying to
         | make it as slow as possible.
        
       | Quitschquat wrote:
       | Tbh, 15% faster than slow AF is still slow AF
        
         | dingdingdang wrote:
         | Yup, but 5 to 15% faster year on year is real progress and
         | that's ultimately what the big user base of Python are counting
         | on at this point.. and they seem to be getting it! Full
         | disclaimer: I'm not a heavy Python user exactly due to the
         | performance and build/distribution situation - it's just sad
         | from a user-end perspective (I'm not addressing centralised web
         | deployment here but rather decentralised distribution which I
         | ultimately find more "real" and rewarding).
        
       | horizion2025 wrote:
       | I don't understand this focus on micro performance details...
       | considering that all of this is about an interpretation approach
       | which is always going to be slow relatively speaking. The big
       | speed up would be to JIT it all, then you dont need to care about
       | structuring of switch loops etc
        
       | eab- wrote:
       | My understanding is that also this tail call based interpretation
       | is also kinder to the branch predictor. I wonder if this explains
       | some of the slow downs - they trigger specific cases that cause
       | lots of branch mispredictions.
        
       | DrewADesign wrote:
       | After years of admonition discouraging me, I'm using Python for a
       | Windows GUI app over my usual C#/MAUI. I'm much more familiar
       | with Python and the whole VS ecosystem is just so heavy for
       | lightweight tasks. I started with tkinter but found it super
       | clunky for interactions I needed heavily, like on field change,
       | but learning QT seemed like more of a lift than I was interested
       | in. (Maybe a skill issue on both fronts?) Grabbed wxglade and
       | drag-and-dropped an interface with wxpython that only has one
       | external dependency installable with pip, is way more convenient
       | than writing xaml by hand, and ergonomically feels pretty
       | pythonic compared to QT. Glad to see more work going into the
       | windows runtime because I'll probably be leaning on it more.
        
         | halfcat wrote:
         | Wait until you see ImGui bindings for Python [1]. It's
         | immediate mode instead of retained mode like Tkinter/Qt/Wx. It
         | might not be what you'd want if you're shipping a thick client
         | to customers, but for internal tooling it's awesome.
         | imgui.text(f"Counter = {counter}")         if
         | imgui.button("increment counter"):             counter += 1
         | _, name = imgui.input_text("Your name?", name)
         | imgui.text(f"Hello {name}!")
         | 
         | [1] https://github.com/pthom/imgui_bundle
        
         | NetMageSCW wrote:
         | Depending on how important the GUI is to you, I would look into
         | LINQPad for stuff that is scripting but too heavy.
        
       | vednig wrote:
       | Python's recent developments have been monumental, new versions
       | now easily best the PyPy performance charts on M4 MacBook Air,
       | idk if this has something to do with optimizations by Apple but
       | coming from Linux I was surprised
        
       | bboreham wrote:
       | Matt Godbolt was saying recently that using tail-calls for an
       | interpreter suits the branch predictor inside the cpu. Compared
       | to a single big switch / computed jump.
        
       | croemer wrote:
       | 2 typos in first sentence. Is this on purpose to make it
       | obviously not-AI generated?
       | 
       | "apology peice" and "tail caling"
        
         | wk_end wrote:
         | If you want to make your writing appear non-AI generated, the
         | easiest way is to write it yourself. No typos necessary.
         | 
         | I'm sure with enough cajoling you can make the LLM spit out a
         | technical blog post that isn't discernibly slop - wanton emoji
         | usage, cliches, self-aggrandizement, relentlessly chipper tone,
         | short "punchy" paragraphs, an absence of depth, "it's not just
         | X--it's a completely new Y" - but it must be at least a little
         | tricky what with how often people don't bother.
         | 
         | [ChatGPT, insert a complaint about how people need to ram LLMs
         | into every discussion no matter how irrelevant here.]
        
         | kenjin4096 wrote:
         | Woops, thanks for noticing, fixed!
        
       | wk_end wrote:
       | So...if the Python team finds tail calls useful, when are we
       | going to see them in Python?
        
       | 01300415352 wrote:
       | Hi
        
       ___________________________________________________________________
       (page generated 2025-12-25 23:00 UTC)