Pandas Should Go Extinct

(eddie.codes)

71 points | by __eddie__ 1 hour ago

20 comments

  • dweinus 1 minute ago
    Nice post but they quickly disregard Dask, don't explain why, don't test it, even exclude it from the benchmark they quote. I don't know if is the better answer, but it seems worth testing if you want a balance approachable + scalable. That's kind the thing Dask was meant to do.
  • minimaxir 1 hour ago
    It's been a while since I've seen an actual data science post submitted to Hacker News: both because AI has superset a lot of DS tasks (e.g. vector embeddings), but also because not much new has happened in DS. Polars has been around for a bit and as noted it is much better than pandas, but otherwise the DS ecosystem has been somewhat stagnant.

    I'd write more tutorials about how to use data science tooling but one consequence of AI is that all the old data sources I used to analyze such as social media and Reddit are now completely locked down (I am surprised NYC Taxi is still being updated, though). Therefore in the meantime, I'm working on making better data science tooling...although unclear to what end due to the data issue above.

  • crazysim 1 hour ago
    Am I crazy or did the OP swap the contents of the posts around accidentally?

    https://eddie.codes/posts/pandas-should-go-extinct/ <=> https://eddie.codes/posts/source-code-comments/

  • sjtrny 1 hour ago
    > People typically start with Excel and graduate to Pandas somewhere in the GB range. Pandas serves them well into the 10s of GBs range, and then they start hitting memory issues, slow computation, or become frustrated with Pandas’ baroque API.

    Assumes that a project moves beyond 10s of GBs. I guess 99.9% of projects that import pandas fall well below this threshold.

    • __eddie__ 1 hour ago
      True, but also if Polars and DuckDB offer a similar experience with the ability to scale beyond that range, why not use them (for new projects)?

      This isn't a call to arms to rewrite everything in the new shiny, just consider the new shiny for new shiny things

      • appplication 48 minutes ago
        Not to mention less footguns. I used to spend my days unwinding bad habits DS pick up from years of panda abuse.

        Credit to pandas for popularizing dataframes in Python, but polars and duckdb are objectively better APIs in addition to their implementation improvements. Agree that it’s time we let it go.

        • __eddie__ 39 minutes ago
          Yes, this was something I emphasised more in the slides / my actual talk. But it's just significantly easier to reason about the APIs for both DuckDB and Polars
    • minimaxir 1 hour ago
      At my work I had convinced the ML pipeline engineers to switch from pandas to polars for even small ETL pipelines and there were notable performance gain with better CPU/memory utilization.

      If a library is performant at large datasets, it is likely performant at small ones too.

      • sjtrny 1 hour ago
        I’m not disagreeing with that statement at all. You missed my point that there are thousands of people making small Python scripts for education and personal projects everyday. In those circumstances the performance concerns are irrelevant and the ergonomics of good pandas documentation and community knowledge make it a better choice.
        • minimaxir 1 hour ago
          That inertia is not a good thing, and it's partially why there's stagnation in data science. Polars is more than mature enough in both documentation and resources for it to be a daily driver.
  • bijowo1676 16 minutes ago
    Strongly disagree with the author.

    For dumb simple select group by OLAP queries on medium data ? Sure clickhouse local or duckdb works perfectly well.

    But if you need to construct dataset ? Or process existing dataset, do heavy filtering, transformation, reshaping, splitting? The proper ETL work, then pandas is really the perfect use case.

    And pandas can work with small memory footprint as well, its actually trivial to do that, plus there are libraries like Modin that are upgrades over pandas with pandas api

    Pandas is the swiss knife tool of data science that lets you do anything with the data and it integrates well with ML libraries

  • akdor1154 28 minutes ago
    Sup Eddie, the actual motivating example is we finally nixed Pandas from Data Ingestion, now reading excel files takes 2 seconds instead of 2 minutes. However unfortunately your

    > TODO: rewrite this entire service

    remains.

    • __eddie__ 23 minutes ago
      > Sup Eddie, the actual motivating example is we finally nixed Pandas from Data Ingestion, now reading excel files takes 2 seconds instead of 2 minutes. However unfortunately your

      I shall pass this information along to my sleep paralysis demons. They'll be glad to hear it.

      > > TODO: rewrite this entire service

      Part of me enjoys the historical significance of this. However, I feel like keeping it would violate my own princples: https://eddie.codes/posts/source-code-comments/

  • japgolly 15 minutes ago
    Nitpick for the author: it looks like you've got (a,b] when you actually mean [a,b). Either that or change the ≥ to be >.
    • __eddie__ 9 minutes ago
      Sorry, where are you seeing this? In one of the code samples?
  • stephenlf 47 minutes ago
    Besides the performance benefits, I use Polars at work because it’s just (subjectively) nicer to work with. The “pl.col” API lets you create arbitrary generated/virtual columns anywhere you want, declaratively. You can throw in these column expressions in wherever without actually computing their values and storing that in memory. Very powerful stuff.
  • willsmith72 1 hour ago
    Makes sense, especially with AI coding tools the rewrite and familiarity arguments hold less water. Similar for the rustify everything crazy.

    The problem is, orgs who see themselves as big data orgs want to act that way, even if they're medium data. "But we'll need it when we grow", "we need to know the state of the art tools"

  • evolve-maz 1 hour ago
    Only in the last few years did I start using SQL properly. Before that my pipelines would live in python. Now I offload as much to the db as possible, and keep my python simple glue. I'm very happy with this compared to other methods in pandas or polars.

    If I still need to do db-like things in python I think duckdb is better.

  • jonahss 1 hour ago
    Real link here: https://eddie.codes/posts/source-code-comments/

    Something going wonky on their blog, where two posts got their links swapped.

    • __eddie__ 1 hour ago
      Yep sorry, fixing now. It's what you get for being bitten by the bug to write twice in the same day
    • jivanvl 1 hour ago
      So it wasn’t just then, for me it linked to their “Useful code comments” post
  • dvt 1 hour ago
    I've been saying this since using Databricks at a company almost a decade ago. Most folks do not need big data tools, and it's just so entrenched because everyone wanted to be a "big data" company and pandas was how you handled big data.
  • stephantul 1 hour ago
    Agreed on all counts.

    In many cases I’ve found directly using python primitives to be less confusing than pandas.

    Similarly, in companies I’ve worked at, the datasets just aren’t that big. Especially if you’ve got access to modern hardware.

  • pjjpo 24 minutes ago
    Liked the title
  • qwertytyyuu 54 minutes ago
    Pandas should go extinct? I’m confused Edit: Oh link was broken before
  • Vaslo 1 hour ago
    I use polars or duckdb now exclusively. Better syntax, better performance. But pandas is deeply entrenched - I try to get my team off it but it’s an uphill battle. It’s not going anywhere anytime soon.
  • jgalt212 58 minutes ago
    I'd drop pandas if polars worked eamlessly with sklearn.
  • ChrisArchitect 1 hour ago
    Title is currently, err...: Useful Code Comments

    Hoping OP can fix this on their end so the url has the expected content. Whoops!

  • viccis 1 hour ago
    Polars seems nice but in my experience using it, the "lazy" APIs would still immediately materialize a ton of stuff in memory and had very spotty support on what data formats and storage integrations were possible with scan_* functions (though that was half a year ago and the support is slowly improving). It's frustrating, I mean really frustrating, to think I could solve a lot of my "scan through heinous amounts of data without any memory hungry things like window aggregations without blowing out my memory" with Polars and then watch my scan_thisorthat() call result in instant memory usage ballooning.

    DuckDB on the other hand is wonderful and truly doesn't use any more memory than it really needs to.

  • jijji 54 minutes ago
    so this story is something that's really important for everybody to know about and should not get downvoted...