[HN Gopher] Why Pandas feels clunky when coming from R (2024)
       ___________________________________________________________________
        
       Why Pandas feels clunky when coming from R (2024)
        
       Author : Tomte
       Score  : 66 points
       Date   : 2025-06-07 16:20 UTC (6 hours ago)
        
 (HTM) web link (www.sumsar.net)
 (TXT) w3m dump (www.sumsar.net)
        
       | great_wubwub wrote:
       | I have no R experience but have been using Polars instead of
       | Pandas for this sort of stuff and it feels less clunky. How does
       | Polars compare to R?
        
         | j_bum wrote:
         | I strongly prefer `dplyr` and the R stack for table processing
         | and visualization.
         | 
         | But, recently I've been working with much larger scale data
         | than R can handle (thanks to R's base int32 limitation) and
         | have been needing to use Python instead.
         | 
         | Polars feels much more intuitive and similar to `dplyr` to me
         | for table processing than Pandas does.
         | 
         | I often ask my LLM of choice to "translate this dplyr call to
         | Polars" as I've been learning the Polars syntax.
        
           | aydyn wrote:
           | It blows my mind that in 2025 R is still limited to 2^31-1
           | rows. R needs a Python 3.0 moment, but that is unfortunately
           | not going to happen for certain unfortunate but unnecessary
           | reasons.
        
             | j_bum wrote:
             | Yep. I have a deep love/hate relationship with R.
             | 
             | This is one of those decisions that I just do not
             | understand. In your mind, why do you imagine a set of
             | improvements won't be made?
             | 
             | Otherwise, for now, working with Python and R using the
             | reticulate package in Quarto is perfect for my needs.
             | 
             | If the Positron IDE could get in-line plot visualization in
             | Quarto documents like the RStudio IDE has, I'd be the
             | happiest camper.
        
               | aydyn wrote:
               | > In your mind, why do you imagine a set of improvements
               | won't be made?
               | 
               | The problem is not technical. Let's just leave it at
               | that.
        
               | j_bum wrote:
               | Ugh now I am extremely curious. This is a lead at least
               | lol.
        
       | BDPW wrote:
       | I've had a similar experience from the opposite side. I've had
       | quite a few years of experience in Python and had to work in R
       | for an internship during my masters.
       | 
       | My impression was that it's pretty easy to do straightforward
       | things like the examples described in the article. But when you
       | have to do complicated or unusual things with your data I found
       | it very frustrating to work with. Access to the underlying data
       | was often opague and it was difficult to me at times to figure
       | out what was happening under the hood.
       | 
       | Does anyone here know any research areas still using R?
        
         | Tomte wrote:
         | Everyone in statistics, and lots of people applying statistics
         | in other disciplines (anthropology etc.).
        
         | j_bum wrote:
         | In addition to stats, R is widely used in computational biology
         | and bioinformatics domains. It's also widely used in the
         | biopharma industry for a variety of other purposes.
        
           | mauritsd wrote:
           | IME (bioinformatics PhD in the netherlands a number of years
           | ago) it's mostly still preferred in a (pre-)clinical context,
           | not so much in academia itself
        
         | kgwgk wrote:
         | > My impression was that it's pretty easy to do straightforward
         | things like the examples described in the article. But when you
         | have to do complicated or unusual things with your data I found
         | it very frustrating to work with.
         | 
         | That's where I realised that the "modern" approach was taken in
         | the article - which obviously I had not looked at.
        
         | pteetor wrote:
         | R is used extensively in quant finance. The quant traders,
         | portfolio managers, and risk managers with whom I work all use
         | R.
        
         | vharuck wrote:
         | As an R user, I get what you mean. If you need to do things
         | that don't fit well in the "tidyverse" model, you have three
         | options:
         | 
         | 1. Wrap the complicated bits in functions, then force it into
         | the tidyverse model by abusing summarize and mutate.
         | 
         | 2. Use data.table. It's very adaptable and handles arbitrary
         | multiline expressions (returning a data.table if the last
         | expression returns a list, otherwise returning the object as-
         | is).
         | 
         | 3. Use base R. It's not as bad as people make it out to be.
         | You'll need to learn it to anyway, if you want to do anything
         | beyond the basics.
        
         | tyfon wrote:
         | Not really research pr se, but it's used extensively in banking
         | here in Norway for anything from statistical model development
         | to basic analysis and reporting.
        
       | dkdcio wrote:
       | pandas* per the style guide (nobody follows it)
       | 
       | also I recommend trying Ibis. created by the creator of pandas
       | originally and solves so many of the issues
       | 
       | https://ibis-project.org
        
         | jna_sh wrote:
         | Any thoughts on ibis vs polars?
        
           | gnulinux wrote:
           | Disclaimer: Never used Ibis before but I daily use polars and
           | DuckDB.
           | 
           | It seems like Ibis uses DuckDB on its backend (by default)
           | and has Polars support as well. Given this, maybe see if Ibis
           | works better for you than polars. If you very specifically
           | need polars, using that will for sure be better. DuckDB is
           | faster than polars and it has great polars support, so
           | depending on how Ibis is implemented it might be "better"
           | than polars as data frame lib.
        
         | Vaslo wrote:
         | There is also Nahwhals.
         | 
         | https://pypi.org/project/narwhals/#description
         | 
         | I tried really hard to use Ibis but I ran into issues where it
         | was way easier to do some stuff in pandas/polars and had to
         | keep coming out of Ibis to make it work so I gave up on it for
         | the time being.
        
       | dleather wrote:
       | I couldn't agree more. I'm fluent in languages like Julia, and
       | MATLAB. I'm 90% fluent in R and prefer data.table over dplyr but
       | working in both is easy enough. The past few months I've been
       | fully transitioning to Python. And while base Python I find to be
       | extremely elegant, typical data science and scientific computing
       | workflows are a headache. There aren't just 1-2 packages to
       | choose from for each use, every package has it's own syntax,
       | keeping track of Pandas Series vs DataFrames is confusion. Want
       | fast differentiable code? Then rewrite everything in numpy in JAX
       | which requires its own tricks.
       | 
       | What Python desperately needs is a coordinated effort for a core
       | data science /scientific computing stack with a unified
       | framework.
       | 
       | In my opinion, if it weren't for Python's extensive use in
       | Industry and package ecosystem, Julia would be the language of
       | choice for nearly all data science and scientific computing uses.
        
         | hatmatrix wrote:
         | > And while base Python I find to be extremely elegant, typical
         | data science and scientific computing workflows are a headache.
         | 
         | That's my impression as well. Going back to the topic of the
         | original post, pandas only partially implements the idioms of
         | the tidyverse so you have to mix in a lot of different forms of
         | syntax (with lambdas to boot) go get things done. Julia is much
         | nicer, but I find myself using PythonCall more often than I'd
         | like.
         | 
         | Scipy was originally supposed to provide the scientific
         | computing stack, but then many offshoots in the direction of
         | pandas / ibis / JAX, etc. happened. I guess that's what you get
         | with a community-based language. MATLAB has its warts but
         | MathWorks does manage to present a coherent stack on that end.
        
       | wodenokoto wrote:
       | A really, really big part of this is thanks to RStudio, which,
       | when you run a line and write a line will peek into memory to see
       | what the columns in your dataframe is and understand the dplyr
       | DSL to help you auto complete what essentially is non-existing
       | variables.
        
       | smabie wrote:
       | The original sin of Pandas is row indices
        
         | hatmatrix wrote:
         | Actually I like that you can use it as a dictionary of tuples
         | (i.e., rows).
        
         | Vaslo wrote:
         | One of the big benefits of polars over pandas is not dealing
         | with the constant index nonsense. Can't tell you all of the
         | issues I had as a beginner with pandas trying to debug silly
         | index errors.
        
       | emehex wrote:
       | I haven't seriously used R in nearly a decade but I _still_ miss
       | (and think about) dpylr and the hadleyverse...
       | 
       | A few years ago I made a package called "redframes" that tried to
       | "solve" all of my frustrations with pandas, make data wrangling
       | feel more like R, while retaining all the best bits of Python...
       | 
       | Alas, it never really took off. For those curious:
       | https://github.com/maxhumber/redframes
        
         | hatmatrix wrote:
         | Hey this looks pretty tidy.
         | 
         | There is so much hype and luck to widespread adoption, you
         | never know with these things.
        
       | __mharrison__ wrote:
       | Lots of readers of my book, Effective Pandas, say it helps them
       | feel more like they are used to with R...
       | 
       | (I've never used R myself, but certainly have some very strong
       | opinions about Pandas after having written 3 books about it.)
        
       ___________________________________________________________________
       (page generated 2025-06-07 23:01 UTC)