[HN Gopher] Why Pandas feels clunky when coming from R (2024)
___________________________________________________________________
Why Pandas feels clunky when coming from R (2024)
Author : Tomte
Score : 66 points
Date : 2025-06-07 16:20 UTC (6 hours ago)
(HTM) web link (www.sumsar.net)
(TXT) w3m dump (www.sumsar.net)
| great_wubwub wrote:
| I have no R experience but have been using Polars instead of
| Pandas for this sort of stuff and it feels less clunky. How does
| Polars compare to R?
| j_bum wrote:
| I strongly prefer `dplyr` and the R stack for table processing
| and visualization.
|
| But, recently I've been working with much larger scale data
| than R can handle (thanks to R's base int32 limitation) and
| have been needing to use Python instead.
|
| Polars feels much more intuitive and similar to `dplyr` to me
| for table processing than Pandas does.
|
| I often ask my LLM of choice to "translate this dplyr call to
| Polars" as I've been learning the Polars syntax.
| aydyn wrote:
| It blows my mind that in 2025 R is still limited to 2^31-1
| rows. R needs a Python 3.0 moment, but that is unfortunately
| not going to happen for certain unfortunate but unnecessary
| reasons.
| j_bum wrote:
| Yep. I have a deep love/hate relationship with R.
|
| This is one of those decisions that I just do not
| understand. In your mind, why do you imagine a set of
| improvements won't be made?
|
| Otherwise, for now, working with Python and R using the
| reticulate package in Quarto is perfect for my needs.
|
| If the Positron IDE could get in-line plot visualization in
| Quarto documents like the RStudio IDE has, I'd be the
| happiest camper.
| aydyn wrote:
| > In your mind, why do you imagine a set of improvements
| won't be made?
|
| The problem is not technical. Let's just leave it at
| that.
| j_bum wrote:
| Ugh now I am extremely curious. This is a lead at least
| lol.
| BDPW wrote:
| I've had a similar experience from the opposite side. I've had
| quite a few years of experience in Python and had to work in R
| for an internship during my masters.
|
| My impression was that it's pretty easy to do straightforward
| things like the examples described in the article. But when you
| have to do complicated or unusual things with your data I found
| it very frustrating to work with. Access to the underlying data
| was often opague and it was difficult to me at times to figure
| out what was happening under the hood.
|
| Does anyone here know any research areas still using R?
| Tomte wrote:
| Everyone in statistics, and lots of people applying statistics
| in other disciplines (anthropology etc.).
| j_bum wrote:
| In addition to stats, R is widely used in computational biology
| and bioinformatics domains. It's also widely used in the
| biopharma industry for a variety of other purposes.
| mauritsd wrote:
| IME (bioinformatics PhD in the netherlands a number of years
| ago) it's mostly still preferred in a (pre-)clinical context,
| not so much in academia itself
| kgwgk wrote:
| > My impression was that it's pretty easy to do straightforward
| things like the examples described in the article. But when you
| have to do complicated or unusual things with your data I found
| it very frustrating to work with.
|
| That's where I realised that the "modern" approach was taken in
| the article - which obviously I had not looked at.
| pteetor wrote:
| R is used extensively in quant finance. The quant traders,
| portfolio managers, and risk managers with whom I work all use
| R.
| vharuck wrote:
| As an R user, I get what you mean. If you need to do things
| that don't fit well in the "tidyverse" model, you have three
| options:
|
| 1. Wrap the complicated bits in functions, then force it into
| the tidyverse model by abusing summarize and mutate.
|
| 2. Use data.table. It's very adaptable and handles arbitrary
| multiline expressions (returning a data.table if the last
| expression returns a list, otherwise returning the object as-
| is).
|
| 3. Use base R. It's not as bad as people make it out to be.
| You'll need to learn it to anyway, if you want to do anything
| beyond the basics.
| tyfon wrote:
| Not really research pr se, but it's used extensively in banking
| here in Norway for anything from statistical model development
| to basic analysis and reporting.
| dkdcio wrote:
| pandas* per the style guide (nobody follows it)
|
| also I recommend trying Ibis. created by the creator of pandas
| originally and solves so many of the issues
|
| https://ibis-project.org
| jna_sh wrote:
| Any thoughts on ibis vs polars?
| gnulinux wrote:
| Disclaimer: Never used Ibis before but I daily use polars and
| DuckDB.
|
| It seems like Ibis uses DuckDB on its backend (by default)
| and has Polars support as well. Given this, maybe see if Ibis
| works better for you than polars. If you very specifically
| need polars, using that will for sure be better. DuckDB is
| faster than polars and it has great polars support, so
| depending on how Ibis is implemented it might be "better"
| than polars as data frame lib.
| Vaslo wrote:
| There is also Nahwhals.
|
| https://pypi.org/project/narwhals/#description
|
| I tried really hard to use Ibis but I ran into issues where it
| was way easier to do some stuff in pandas/polars and had to
| keep coming out of Ibis to make it work so I gave up on it for
| the time being.
| dleather wrote:
| I couldn't agree more. I'm fluent in languages like Julia, and
| MATLAB. I'm 90% fluent in R and prefer data.table over dplyr but
| working in both is easy enough. The past few months I've been
| fully transitioning to Python. And while base Python I find to be
| extremely elegant, typical data science and scientific computing
| workflows are a headache. There aren't just 1-2 packages to
| choose from for each use, every package has it's own syntax,
| keeping track of Pandas Series vs DataFrames is confusion. Want
| fast differentiable code? Then rewrite everything in numpy in JAX
| which requires its own tricks.
|
| What Python desperately needs is a coordinated effort for a core
| data science /scientific computing stack with a unified
| framework.
|
| In my opinion, if it weren't for Python's extensive use in
| Industry and package ecosystem, Julia would be the language of
| choice for nearly all data science and scientific computing uses.
| hatmatrix wrote:
| > And while base Python I find to be extremely elegant, typical
| data science and scientific computing workflows are a headache.
|
| That's my impression as well. Going back to the topic of the
| original post, pandas only partially implements the idioms of
| the tidyverse so you have to mix in a lot of different forms of
| syntax (with lambdas to boot) go get things done. Julia is much
| nicer, but I find myself using PythonCall more often than I'd
| like.
|
| Scipy was originally supposed to provide the scientific
| computing stack, but then many offshoots in the direction of
| pandas / ibis / JAX, etc. happened. I guess that's what you get
| with a community-based language. MATLAB has its warts but
| MathWorks does manage to present a coherent stack on that end.
| wodenokoto wrote:
| A really, really big part of this is thanks to RStudio, which,
| when you run a line and write a line will peek into memory to see
| what the columns in your dataframe is and understand the dplyr
| DSL to help you auto complete what essentially is non-existing
| variables.
| smabie wrote:
| The original sin of Pandas is row indices
| hatmatrix wrote:
| Actually I like that you can use it as a dictionary of tuples
| (i.e., rows).
| Vaslo wrote:
| One of the big benefits of polars over pandas is not dealing
| with the constant index nonsense. Can't tell you all of the
| issues I had as a beginner with pandas trying to debug silly
| index errors.
| emehex wrote:
| I haven't seriously used R in nearly a decade but I _still_ miss
| (and think about) dpylr and the hadleyverse...
|
| A few years ago I made a package called "redframes" that tried to
| "solve" all of my frustrations with pandas, make data wrangling
| feel more like R, while retaining all the best bits of Python...
|
| Alas, it never really took off. For those curious:
| https://github.com/maxhumber/redframes
| hatmatrix wrote:
| Hey this looks pretty tidy.
|
| There is so much hype and luck to widespread adoption, you
| never know with these things.
| __mharrison__ wrote:
| Lots of readers of my book, Effective Pandas, say it helps them
| feel more like they are used to with R...
|
| (I've never used R myself, but certainly have some very strong
| opinions about Pandas after having written 3 books about it.)
___________________________________________________________________
(page generated 2025-06-07 23:01 UTC)