[HN Gopher] Overhead of Python asyncio tasks
       ___________________________________________________________________
        
       Overhead of Python asyncio tasks
        
       Author : willm
       Score  : 94 points
       Date   : 2023-03-08 18:50 UTC (4 hours ago)
        
 (HTM) web link (textual.textualize.io)
 (TXT) w3m dump (textual.textualize.io)
        
       | masklinn wrote:
       | Seems pretty decent, with that it's a real shame Python's async
       | ended up coroutine-based rather than task-based. Given the
       | langage semantics it ends up being a lot of pain for fairly
       | little gain at the end if the day.
        
       | jstx1 wrote:
       | Async Python is still confusing af - when do I need it, what
       | happens under the hood, does it actually help with performance,
       | sometimes the GIL comes into play and sometimes it doesn't, why
       | do we ever use threads at all if there's a GIL, why is it called
       | async _io_ if we can use it for anything. My mind is kind of
       | scattered and people seems to be using a lot of async Python for
       | some reason.
       | 
       | Any good resources to clear things up?
        
         | uniqueuid wrote:
         | I agree that especially within the standard context of python
         | and its syntax, async seems weird (because it's sprinkled into
         | an existing paradigm).
         | 
         | The best mental model for me always was to think: Here's an
         | await, that means "interpreter, go and do something else that's
         | currently waiting while this is not done yet."
         | 
         | And that's all about IO, because what you can wait on is
         | essentially IO.
         | 
         | By the way, I really wish there was a better story to executors
         | in async python. To think we still have the same queue/pickle-
         | based multiprocessing to send async tasks to another core is
         | kind of sad. Hoping for 3.12 and beyond there.
         | 
         | [edit] one really neat example that helped me get asyncio was
         | Guido van Rossum's crawler in 500 lines of python [1]. A lot of
         | the syntax is deprecated now, but it's still a great walk-
         | through
         | 
         | [1] http://aosabook.org/en/500L/a-web-crawler-with-asyncio-
         | corou...
        
         | ActorNightly wrote:
         | Async looks like parallelism, but its just smart scheduling.
         | The core concept of async is saying "hey, im waiting for
         | something to complete thats not under my control (like waiting
         | for data on a socket to be able to be read), go ahead and do
         | other things in the mean time".
         | 
         | If you ever coded in sockets in C (and its a good exercise to
         | do so), you probably have at some point ran across `select`
         | which is essentially a non blocking way to check which sockets
         | have data available to read, and then sequentially read the
         | data. This gives the ability for a program to appear
         | parallelized in the sense that it can handle multiple client
         | connections, but its not truly parallel. Different clients can
         | be handled at different time depending on which order they
         | connect, which is asynchronous in nature (versus processing
         | each client in sequence and waiting on each one to connect and
         | disconnect before moving on to the next one)
         | 
         | Async in Python is basically this concept, with a core
         | fundamental feature of time limited execution. Functions can
         | say that they are pausing for x seconds, allowing other
         | functions to run, or functions can say that they give a certain
         | function x seconds to run before resuming execution. If you
         | async code (along with any library you may use) doesn't contain
         | any sleeps or timeouts, its exactly equivalent to synchronous
         | code (since the event loop never really recieves a message that
         | it can suspend a routine or cancel it). With sleeps and
         | timeouts, you gain control over things that can potentially
         | block, both from a caller perspective of not having a function
         | call block your own, and from a callee perspective of not
         | making your function blocking.
         | 
         | The use case is for it is that it is good for I/O bound
         | operations like Threading is, but with the addition that you
         | don't have to worry about synchronization or race conditions,
         | since by design your code will have predictable access
         | patterns. The downside is that your code and any libraries that
         | you use within your code has to be implemented as async
         | libraries, and any library that is async has to have async
         | wrappers around the calls to its methods, which in turn means
         | that your entire code has to be async.
         | 
         | Threading with Python is generally not useful, as its not true
         | parallelism because of GIL. GIL allows only one thread in
         | Python to run. Threading is safer in Python because of this,
         | however it obviously has drawbacks. In general its best used if
         | you want asyncio like performance with a library that is not
         | written with async, since GIL is smart enough to detect when a
         | thread is waiting for input and switch context.
         | 
         | True parallelism in Python is achieved with multiprocessing,
         | however the use case is a little different. Rather than
         | spinning off processes, you generally launch a bunch of worker
         | processes up front (to avoid the larger overhead), then use
         | smart scheduling to distribute work between these processes.
         | Here though you do have to worry about race conditions and
         | synchronization, and use things like locks and mutexes.
        
         | bombolo wrote:
         | It's basically a fancy way to do epoll() around file
         | descriptors, but hide the need to keep a main loop and state,
         | and keep it hidden as functions that run in fake concurrency,
         | stopping whenever they block, and being executed again when
         | their file descriptor has activity.
         | 
         | It doesn't necessarily improve performance. It's just a much
         | easier way to do non-blocking I/O (note that blocking and
         | threaded I/O is easier to do but much much heavier).
        
         | 323 wrote:
         | It makes concurrent programming much simpler than using
         | threads.
         | 
         | Very few locking and care is needed with asyncio, as opposed to
         | using threads. Race conditions are basically not a thing if you
         | write reasonably idiomatic code.
         | 
         | It might be (or not) faster than using threads, but that's not
         | the main benefit in my view - this easiness of use is.
        
         | scrlk wrote:
         | I've found SuperFastPython to be helpful for understanding
         | Python concurrency: https://superfastpython.com/python-
         | concurrency-choose-api/
        
         | philote wrote:
         | It's great for when you can do concurrent I/O tasks. For
         | example running a web backend, web scraping, or many API calls.
         | FastAPI (and Starlette that it's built off of) is an async web
         | framework and in my experience performs well.
         | 
         | Basically, your program will normally halt when doing I/O, and
         | won't proceed until that I/O is done. During that halt, your
         | program is doing nothing (no cpu being used). With asyncio, you
         | can schedule multiple tasks to run, so if one is halted doing
         | I/O, another can run.
         | 
         | Edit: And AFAIK, the GIL does not come into play at all with
         | async. Only when multithreading.
        
           | morelisp wrote:
           | > the GIL does not come into play at all with async.
           | 
           | In the sense that the GIL is still held and you can have at
           | most one path of Python code executing at a time regardless
           | of whether you use async or threads, sure. Most blocking I/O
           | was already releasing the GIL so the difference is purely in
           | how you can design your modules; for any reasoning about
           | performance the GIL behaves the same way whether you use
           | asyncio or not.
        
       | Hendrikto wrote:
       | > Clearly create_task is as close as you get to free in the
       | Python world, and I would need to look elsewhere for
       | optimizations. Turns out Textual spends far more time processing
       | CSS rules than creating tasks (obvious in retrospect).
       | 
       | Takeaways:
       | 
       | 1. Creating async tasks is cheap. 2. It is important to confirm
       | intuitions, before acting on them.
        
       | TickleSteve wrote:
       | So, 250,000 per second on a (roughly 10,000 MIPS I7 core)
       | 
       | Each of those task_create calls is roughly 10,000,000,000 /
       | 250,000 = 40,000 instructions.
       | 
       | Thats 40'000 instructions of pure overhead as it does not
       | contribute to the task at hand (accidental complexity).
        
         | schmichael wrote:
         | Your method of estimating instructions includes printing
         | multiple lines which context switch to the kernel to perform
         | IO. I'm not sure how an "instruction count" metric is useful
         | anyway.
         | 
         | edit: I'm not actually sure you are counting the context
         | switch, but I still don't think estimating instruction count
         | that way is particularly useful.
        
           | TickleSteve wrote:
           | Its useful to show how much waste there is in a solution.
           | 
           | Having an operation that can be executing 250,000 times per
           | second on a modern processor is extremely _slow_... not fast.
        
             | 323 wrote:
             | Waste allows scale, that's a general civilizational rule.
             | 
             | You wouldn't be able to write your comment if the browser
             | were written in extremely efficient assembly code because
             | it wouldn't exist.
             | 
             | NASA and every big organization also has a lot of waste,
             | but only that way you can get to the moon.
        
             | schmichael wrote:
             | 250k/s is roughly the same speed as context switching, so
             | while slow for pure computation, it is a reasonable amount
             | of "waste" for switching between concurrent tasks.
             | 
             | If you didn't prevent preemptive context switches during
             | your benchmarking, it's entirely possible the only thing
             | you measured was the context switch time.
             | 
             | This is a fun experiment, but to get a rigorous idea of the
             | overhead involved takes more work than what anyone in the
             | post or comments has done.
        
               | morelisp wrote:
               | Matching the cost of a genuine context switch should be a
               | (laughably bad) _upper_ bound for any language 's
               | particular concurrency offerings. It is not reasonable.
        
               | schmichael wrote:
               | > It is not reasonable.
               | 
               | Reasonableness is relative and use case dependent. The
               | post itself illustrates how the cost is insignificant
               | compared to other "wasteful" operations related to CSS
               | handling.
               | 
               | If this is too much overhead for your use case, there are
               | plenty of other approaches and languages to choose from.
        
               | morelisp wrote:
               | If it costs as much as a context switch, you might as
               | well just do context switches. These hosted language
               | scaffoldings - whether that's asyncio, go routines, TPL,
               | Webflux, etc. - exist specifically so you don't have to
               | do a full context switch. If they cost as much as a
               | context switch, they have failed. Regardless of what else
               | is taking time in the system.
               | 
               | If you're not any better, just replace your whole hosted
               | concurrency system with a statement that triggers
               | sched_yield.
        
       | mkl95 wrote:
       | Python is my strongest language. If you ask me to write some
       | asynchronous code I will try my best not to write it in Python.
       | Usually it's just not worth it.
        
         | philote wrote:
         | Have you not used async in Python lately? IMO it's very easy to
         | do.
        
       | t3pfaff wrote:
       | Why would you want a terminal emulator anywhere near python?
       | Using python for lightweight system utility gui apps seems like
       | using a hammer to screw in a nail. Yeah you can do it and modern
       | hardware is fast enough that you probably won't care, but why??
        
         | traverseda wrote:
         | Python is a really good scripting language? I guess the big
         | alternatives would be C and BASH and I'd pick python for short
         | utilities over both of those any day.
         | 
         | How slow do you think python is?
        
           | vishvananda wrote:
           | Python is extremely slow for some tasks. I was surprised to
           | discover how slow when I ran some benchmarks, despite having
           | used python for many years at the time. It has been improving
           | lately, but here is a blog post I made on the topic quite a
           | few years ago that has some interesting comparisons:
           | https://gist.github.com/vishvananda/7a2f1942d0e9ffff4093
        
             | vishvananda wrote:
             | Just reran the benchmarks from 10 years ago, python is only
             | 37X slower than C on the benchmark now, and the go version
             | is running faster than the C version. Python still has big
             | productivity wins of course...
        
           | t3pfaff wrote:
           | I agree that it is a good scripting language. That's where it
           | excells. A GUI terminal emulator however, is not a script...
           | Python once compiled and using it's underlying c libraries is
           | fast enough by modern standards but it is slow to start and
           | projects using it are prone towards difficult to read/messy
           | code.
           | 
           | Obviously the latter is up to the developer(s) and whatever
           | standards they set for themselves but any software project
           | tends towards paths of least resistance inherent in the
           | language and frameworks being used over time as maintainers
           | change and PRs fixes from contributers are merged. This is
           | pedantic of course, but that doesn't change the fact that
           | Python is not the optimal solution for this problem.
        
             | simonw wrote:
             | I think you're misunderstanding what Textual is. It's not a
             | GUI terminal emulator: it's a library for building
             | interactive terminal applications.
        
               | t3pfaff wrote:
               | Yeah I realized that later.
        
             | mixmastamyk wrote:
             | You seem to be off by an order or three of magnitude about
             | what is "fast enough" for terminal apps, which haven't been
             | a performance bottleneck for decades.
             | 
             | 200k ops per second not fast enough for you? How many words
             | per second do you type?
             | 
             | So why wouldn't you use it? I've have since the late 90s
             | and even then it was faster than I could keep up with.
             | Unoptimized Java Swing on SGI was the only thing that
             | wasn't, from memory. More recently Windows Terminal had a
             | very bad implementation for a while, but fixed it after
             | their ass was handed to them over it here at HN.
        
               | t3pfaff wrote:
               | I really don't know why you are so focused on the speed.
               | Obviously python merely acting as translation to
               | optimized c libraries is going to be fast enough. I said
               | the performance difference was negligible on modern
               | hardware. Python's speed is far from its largest problem
               | as I outlined in the above comment.
        
               | mixmastamyk wrote:
               | I see, thought it was the main point, but there were
               | others...
               | 
               | > but it is slow to start
               | 
               | This is not really true either. Sure not as fast as C,
               | but imperceptible for the most part. And they have
               | improved it in recent versions.
               | 
               | Yes, you can _cause_ it be an issue with a poorly written
               | or though out system, this happened at one job I had. But
               | that wasn 't Python's fault they decided to pull in
               | thousands of files each invocation.
               | 
               | My Python scripts respond instantly, even big ones. I
               | have a CLI photo editor and implementation lang is not an
               | issue that I even contemplated until now.
               | 
               | > and projects using it are prone towards difficult to
               | read/messy code.
               | 
               | Primarily large ones with a long history of alternating
               | developers. There are great tools to improve its
               | scalability; use them. The simple pyflakes will eliminate
               | most issues. Type checking gets the long tail for the
               | mission-important+.
        
         | [deleted]
        
         | throwaway23422_ wrote:
         | [flagged]
        
       | l_theanine wrote:
       | I haven't used asyncio that much, certainly not in any serious
       | sense, but wouldn't ContextVar lookups be a major factor of
       | performance in serious asyncio code? Using tasks for things that
       | aren't io-bound seems likely to give a false sense of performance
       | superiority when doing basically nothing.
        
         | willm wrote:
         | I'd be surprised if context vars are more expensive than a dict
         | lookup or two. But I haven't profiled. Could be wrong.
        
       | Kab1r wrote:
       | I have been looking into the overhead if async on c++ and we
       | found that the cost increases substantially when the function has
       | a return value. It would be interesting to see if this is the
       | case with python 's asyncio.
        
       | uniqueuid wrote:
       | async tasks are cool, but the usual PSA applies here:
       | 
       | Be careful to hold your references, because async tasks without
       | active references will be garbage collected. I've been bitten by
       | that in the past.
       | 
       | Long discussion here: https://bugs.python.org/issue21163
       | 
       | Docs: https://docs.python.org/3/library/asyncio-
       | task.html#asyncio....
       | 
       | "Important
       | 
       | Save a reference to the result of this function, to avoid a task
       | disappearing mid-execution. The event loop only keeps weak
       | references to tasks. A task that isn't referenced elsewhere may
       | get garbage collected at any time, even before it's done."
        
         | jonathan_s wrote:
         | If you can, best is to always spawn them in a task group
         | (either using anyio or Python 3.11's task groups).
         | 
         | This prevents tasks from being garbage collected, but also
         | prevents situations where components can create tasks that
         | outlive their own lifetime. Plus, it's a saner approach when
         | dealing with exception handling and cancellation.
        
           | uniqueuid wrote:
           | Perhaps I just don't get them, but task groups never really
           | made sense to me.
           | 
           | The whole beauty of async tasks is that you can spawn, retry,
           | and consume them lazily. When you create a task group, you
           | again end up waiting on a single long-running last task,
           | desperately trying to fix individual failures and retries
           | that hold up the entire group.
        
         | gmadsen wrote:
         | task groups solve this problem
        
         | mrnonchalant wrote:
         | Thanks for this!
        
       | tpmx wrote:
       | Mostly a testament to how absurdly fast modern CPUs are despite
       | Python itself being so slow.
       | 
       | / Recent Python convert, in spite of the horrible general
       | performance of the official implementation of the language. That
       | sweet, sweet module library. Also, with Docker containers the
       | deployment issues have been solved. It might be slow to execute
       | but it's really efficient to develop with.
        
       | chinaman425 wrote:
       | [dead]
        
       | crabbone wrote:
       | Reading this is like reading early Renaissance alchemist arguing
       | about how much mercury they need to combine with how much silver
       | to create gold... This is so far gone I don't even know where to
       | begin...
       | 
       | > It may be IO that gives AsyncIO its name, but Textual doesn't
       | do any IO of its own.
       | 
       | So why on Earth are you using AsyncIO? You don't need it, if
       | that's true...
       | 
       | > Those tasks are used to power message queues
       | 
       | How are your message queues not doing I/O? What on Earth are they
       | doing then?
       | 
       | Needless to say that the whole benchmark is worthless because it
       | never even initiates anything that would be involved when
       | creating actual asynchronous I/O tasks...
       | 
       | ----
       | 
       | I mean, I know, in Pythonland this is just your average
       | Wednesday, but dear lord, if you don't visit that land all that
       | often it shocks you more every time you do.
        
         | louwrentius wrote:
         | I knew you would be here in the comments. Crabbone won't leave
         | any chance he gets to shit and dump on Python any HN post he
         | can.
         | 
         | Must be so frustrating, knowing that this silly language is one
         | of the most populair programming languages in the world,
         | despite its short comings, one of which is slower performance.
         | 
         | This isn't about python or any language. The extreme negative
         | tone and sense of superiority seems to point to something else.
         | 
         | What is it?
        
         | vore wrote:
         | asyncio is really not all about I/O though, despite the name.
         | You can just as easily use it for concurrency by interleaving
         | tasks with each other. You can imagine multiple UI elements
         | that all need to make progress, and using coroutines to have
         | them cooperatively yield to each other and not just have one
         | element block progress everywhere else.
         | 
         | It's just using asyncio as a task scheduler, nothing more,
         | nothing less. Maybe when you visit Pythonland you might instead
         | be shocked in a more positive way :-)
        
         | greenshackle2 wrote:
         | Asyncio is for cooperative multitasking. I/O is the most common
         | use case but it's not the only one. They're using the event
         | loop to schedule their GUI tasks.
         | 
         | Textual is a framework for building desktop apps.
         | 
         | I assume the message queues are in-memory structures used to
         | pass messages between tasks, hence no I/O.
         | 
         | This is really fairly standard stuff.
         | 
         | I understand you may not be familiar with GUI software and/or
         | the Python ecosystem but jumping straight to condescension when
         | you don't understand something is not really a good attitude.
        
         | Uptrenda wrote:
         | This reply strongly reads of 'im smarter than eveyone' and
         | 'know everything.' Is it possible that the author is doing
         | something that you don't know about? Such as writing and
         | reading text to sockets (which would be I/O.) And is it
         | possible they may have a reason to be using asyncio in Python
         | (which uses a single main thread and hence needs something like
         | asyncio for concurrency.)
         | 
         | >Needless to say that the whole benchmark is worthless
         | 
         | Overhead startup isn't worthless.
         | 
         | 'I mean, I know, in Pythonland this is just your average
         | Wednesday'
         | 
         | and here we go. The whole post was really just you trying to
         | make out that you're superior to everyone else. In this case
         | looking down on an entire ecosystem. Why? Python is one of the
         | most popular languages. It has an elegant syntax and it's
         | capable of solving most problems. Go jerk yourself off in
         | private.
        
         | traverseda wrote:
         | Python asyncio is actually a co-operative multitasking system,
         | think real-time operating system type stuff. It doesn't do real
         | concurrency but it gives you a nice interface for reasoning
         | about tasks (think thread) while being able to explicitly
         | control when/where they yield control of the event loop, how
         | often they run, etc.
         | 
         | >How are your message queues not doing I/O?
         | 
         | We generally wouldn't consider in-process moving data around to
         | be "I/O", now if you started interacting with an external
         | database/file/pipe/MMAP-ed-file/etc than that would be "I/O".
         | 
         | >What on Earth are they doing then?
         | 
         | Tasks. Kind of like processes but lighter weight. Running an
         | event loop, dispatching signals, that sort of thing. You can do
         | all that by manually writing your own event loop but python's
         | asyncio (IMO) makes it easier to reason about exactly when your
         | yielding the event loop to some other task, and makes it easier
         | to write code that can yield control of the event loop at
         | arbitrary places. So like if you want to update a widget every
         | 10 seconds you can write something like                   while
         | True:             await asyncio.sleep(10) #Other tasks can run
         | during this 10 seconds sleep             #Update widget
         | contents
         | 
         | Useful for stuff that needs to periodically poll data, like a
         | process monitor, or even for just simple clock widgets.
         | 
         | ---
         | 
         | As an aside that cooperative multitasking can be really nice in
         | micropython, where you can write tight loops in straight
         | assembly if you need to and still get a pleasant task interface
         | for managing higher level tasks/threads. Combined with some
         | interrupt handlers it makes a pretty elegant real-time-ish
         | operating system. (You probably need to manually deal with
         | garbage collections though)
        
         | philote wrote:
         | Message queues can be implemented in memory, so no network or
         | disk I/O. In fact, it seems this project does use the built-in
         | "queue" library: https://docs.python.org/3/library/queue.html
        
       | johndubchak wrote:
       | After running that code on both a Windows SB3 and major souped up
       | Lenovo running Ubuntu...I just feel inadequate.
        
         | jayde2767 wrote:
         | Ubuntu:                 100,000 tasks   155,257 tasks per/s
         | 200,000 tasks   138,569 tasks per/s       300,000 tasks
         | 134,779 tasks per/s       400,000 tasks   144,371 tasks per/s
         | 500,000 tasks   135,672 tasks per/s       600,000 tasks
         | 135,299 tasks per/s       700,000 tasks   146,456 tasks per/s
         | 800,000 tasks   139,192 tasks per/s
        
         | SCUSKU wrote:
         | M2 Macbook Pro 16GB 16-inch 2023
         | 
         | 100,000 tasks 184,167 tasks per/s
         | 
         | 200,000 tasks 160,964 tasks per/s
         | 
         | 300,000 tasks 165,278 tasks per/s
         | 
         | 400,000 tasks 149,577 tasks per/s
         | 
         | 500,000 tasks 160,593 tasks per/s
         | 
         | 600,000 tasks 168,098 tasks per/s
         | 
         | 700,000 tasks 161,837 tasks per/s
         | 
         | 800,000 tasks 160,364 tasks per/s
         | 
         | 900,000 tasks 149,479 tasks per/s
         | 
         | 1,000,000 tasks 155,919 tasks per/s
        
         | zomnoys wrote:
         | Above anything, this shows the performance gains from 3.10 ->
         | 3.11:                 >> python3.10 create_task_overhead.py
         | 100,000 tasks   185,694 tasks per/s       200,000 tasks
         | 165,581 tasks per/s       300,000 tasks   170,857 tasks per/s
         | 400,000 tasks   159,081 tasks per/s       500,000 tasks
         | 162,640 tasks per/s       600,000 tasks   158,779 tasks per/s
         | 700,000 tasks   161,779 tasks per/s       800,000 tasks
         | 179,965 tasks per/s       900,000 tasks   160,913 tasks per/s
         | 1,000,000 tasks  162,767 tasks per/s            >> python3.11
         | create_task_overhead.py       100,000 tasks   289,318 tasks
         | per/s       200,000 tasks   265,293 tasks per/s       300,000
         | tasks   266,011 tasks per/s       400,000 tasks   259,821 tasks
         | per/s       500,000 tasks   251,819 tasks per/s       600,000
         | tasks   267,441 tasks per/s       700,000 tasks   251,789 tasks
         | per/s       800,000 tasks   254,303 tasks per/s       900,000
         | tasks   249,894 tasks per/s       1,000,000 tasks  266,581
         | tasks per/s
        
         | nathanasmith wrote:
         | Python 3.11 running in a-Shell on an M1 iPad Pro
         | Python 3.11.0 (heads/3.11-dirty:8d3dd5b9647, Dec  7 2022,
         | 08:17:48) [Clang 14.0.0 (clang-1400.0.29.202)]      on darwin
         | Type "help", "copyright", "credits" or "license" for more
         | information.         >>>          [~/Documents]$ python test.py
         | 100,000 tasks    127,992 tasks per/s         200,000 tasks
         | 115,960 tasks per/s         300,000 tasks    117,205 tasks
         | per/s         400,000 tasks    113,131 tasks per/s
         | 500,000 tasks    109,609 tasks per/s         600,000 tasks
         | 116,649 tasks per/s         700,000 tasks    110,743 tasks
         | per/s         800,000 tasks    111,361 tasks per/s
         | 900,000 tasks    109,688 tasks per/s         1,000,000 tasks
         | 117,064 tasks per/s
        
         | johndubchak wrote:
         | Windows SB3:                 100,000 tasks    177,778 tasks
         | per/s       200,000 tasks    150,588 tasks per/s       300,000
         | tasks    152,381 tasks per/s       400,000 tasks    134,031
         | tasks per/s       500,000 tasks    160,804 tasks per/s
         | 600,000 tasks    129,293 tasks per/s
        
       | brrrrrm wrote:
       | JavaScript is 10-37x faster out of the box without any imports
       | async function time_tasks(count=100) {           async function
       | nop_task() {             return performance.now();           }
       | const start = performance.now()           let tasks =
       | Array(count).map(nop_task)           await Promise.all(tasks)
       | const elapsed = performance.now() - start           return
       | elapsed / 1e3         }              for (let count = 100000;
       | count < 1000000 + 1; count += 100000) {           const ct =
       | await time_tasks(count)           console.log(`${count}: ${1 /
       | (ct / count)} tasks/sec`)         }
       | 
       | Outputs (Python 3.11, Bun 0.5.1):                   % bun
       | textual.ts         100000: 3767797.000743159 tasks/sec
       | 200000: 9001406.4697609 tasks/sec         300000:
       | 8281002.001242148 tasks/sec         400000: 10038491.340232708
       | tasks/sec         500000: 8976653.913474608 tasks/sec
       | 600000: 10437550.828698047 tasks/sec         700000:
       | 9443895.154523576 tasks/sec         800000: 11021991.118011119
       | tasks/sec         900000: 9790550.215324111 tasks/sec
       | 1000000: 10263937.143648934 tasks/sec                  % python3
       | textual.py         100,000 tasks   303,063 tasks per/s
       | 200,000 tasks   270,058 tasks per/s         300,000 tasks
       | 271,621 tasks per/s         400,000 tasks   261,945 tasks per/s
       | 500,000 tasks   251,070 tasks per/s         600,000 tasks
       | 272,520 tasks per/s         700,000 tasks   250,977 tasks per/s
       | 800,000 tasks   253,131 tasks per/s         900,000 tasks
       | 244,696 tasks per/s         1,000,000 tasks   266,061 tasks per/s
        
         | seanw444 wrote:
         | Bun* is that much faster. Wonder what Node or Deno perform
         | like.
        
       | schmichael wrote:
       | What a fun experiment. I quick converted it to Go using
       | goroutines and waitgroups for fun:
       | https://gist.github.com/schmichael/1a417808b8e88b684838ae9f4...
        
         | bob1029 wrote:
         | I am seeing 500~600k tasks/second in .NET6 w/ TPL.
         | 
         | Edit: Updating per request below.
         | 
         | Tested on a TR2950x. I wonder if NUMA issues in my case or bad
         | python version (3.9.4).
         | 
         | Python                 100,000 tasks    77,108 tasks per/s
         | 200,000 tasks    69,945 tasks per/s       300,000 tasks
         | 72,453 tasks per/s       400,000 tasks    74,636 tasks per/s
         | 500,000 tasks    66,253 tasks per/s       600,000 tasks
         | 77,576 tasks per/s       700,000 tasks    69,673 tasks per/s
         | 800,000 tasks    68,176 tasks per/s       900,000 tasks
         | 73,846 tasks per/s       1,000,000 tasks          68,013 tasks
         | per/s
         | 
         | .NET 6                 100000 Tasks 523000 Tasks/s       200000
         | Tasks 550000 Tasks/s       300000 Tasks 550000 Tasks/s
         | 400000 Tasks 559000 Tasks/s       500000 Tasks 547000 Tasks/s
         | 600000 Tasks 539000 Tasks/s       700000 Tasks 547000 Tasks/s
         | 800000 Tasks 540000 Tasks/s       900000 Tasks 560000 Tasks/s
         | 1000000 Tasks 542000 Tasks/s
        
           | lunfard000 wrote:
           | would you mind sharing the go and python results running on
           | your machine too? It is apples to orange comparation
           | otherwise.
           | 
           | EDIT. My results on a 5950x (undervolted)
           | 
           | python3.8.exe test.py
           | 
           | 100,000 tasks 139,130 tasks per/s
           | 
           | 200,000 tasks 121,905 tasks per/s
           | 
           | 300,000 tasks 120,000 tasks per/s
           | 
           | 400,000 tasks 114,286 tasks per/s
           | 
           | 500,000 tasks 119,403 tasks per/s
           | 
           | 600,000 tasks 117,073 tasks per/s
           | 
           | 700,000 tasks 130,612 tasks per/s
           | 
           | 800,000 tasks 122,488 tasks per/s
           | 
           | 900,000 tasks 120,000 tasks per/s
           | 
           | 1,000,000 tasks 110,155 tasks per/s
           | 
           | python3.11.exe .\test.py
           | 
           | 100,000 tasks 206,452 tasks per/s
           | 
           | 200,000 tasks 185,507 tasks per/s
           | 
           | 300,000 tasks 186,408 tasks per/s
           | 
           | 400,000 tasks 179,021 tasks per/s
           | 
           | 500,000 tasks 167,539 tasks per/s
           | 
           | 600,000 tasks 177,778 tasks per/s
           | 
           | 700,000 tasks 188,235 tasks per/s
           | 
           | 800,000 tasks 180,919 tasks per/s
           | 
           | 900,000 tasks 168,421 tasks per/s
           | 
           | .\test.exe (go 1.20 compiled)
           | 
           | 100000 tasks 2710563.336378 tasks per/s
           | 
           | 200000 tasks 3076885.207567 tasks per/s
           | 
           | 300000 tasks 3332292.917434 tasks per/s
           | 
           | 400000 tasks 3040479.422795 tasks per/s
           | 
           | 500000 tasks 2810232.844653 tasks per/s
           | 
           | 600000 tasks 3004138.200371 tasks per/s
           | 
           | 700000 tasks 2738877.029117 tasks per/s
           | 
           | 800000 tasks 2893730.985022 tasks per/s
           | 
           | 900000 tasks 3043877.494077 tasks per/s
           | 
           | 1000000 tasks 2857992.089078 tasks per/s
        
           | chinaman425 wrote:
           | [dead]
        
         | benhoyt wrote:
         | Yeah, I did the exact same thing and got very similar results.
         | Just to save people clicking through, the Go/goroutine version
         | is 25x as fast as Python. :-)
        
       ___________________________________________________________________
       (page generated 2023-03-08 23:01 UTC)