https://github.com/machow/siuba Skip to content Sign up * Why GitHub? Features - + Mobile - + Actions - + Codespaces - + Packages - + Security - + Code review - + Project management - + Integrations - + GitHub Sponsors - + Customer stories - + Security - * Team * Enterprise * Explore + Explore GitHub - Learn & contribute + Topics - + Collections - + Trending - + Learning Lab - + Open source guides - Connect with others + The ReadME Project - + Events - + Community forum - + GitHub Education - + GitHub Stars program - * Marketplace * Pricing Plans - + Compare plans - + Contact Sales - + Nonprofit - + Education - [ ] [search-key] * # In this repository All GitHub | Jump to | * No suggested jump to results * # In this repository All GitHub | Jump to | * # In this user All GitHub | Jump to | * # In this repository All GitHub | Jump to | Sign in Sign up {{ message }} machow / siuba * Watch 7 * Star 480 * Fork 18 Python library for using dplyr like syntax with pandas and SQL siuba.readthedocs.io MIT License 480 stars 18 forks Star Watch * Code * Issues 78 * Pull requests 5 * Actions * Projects 7 * Wiki * Security * Insights More * Code * Issues * Pull requests * Actions * Projects * Wiki * Security * Insights master 13 branches 18 tags Go to file Code Clone HTTPS GitHub CLI [https://github.com/m] Use Git or checkout with SVN using the web URL. [gh repo clone machow] Work fast with our official CLI. Learn more. * Open with GitHub Desktop * Download ZIP Launching GitHub Desktop If nothing happens, download GitHub Desktop and try again. Go back Launching GitHub Desktop If nothing happens, download GitHub Desktop and try again. Go back Launching Xcode If nothing happens, download Xcode and try again. Go back Launching Visual Studio If nothing happens, download the GitHub extension for Visual Studio and try again. Go back Latest commit @machow machow Merge pull request #245 from machow/feat-pipe-compose ... 030eead Jan 8, 2021 Merge pull request #245 from machow/feat-pipe-compose fix: pipe composition with symbols 030eead Git stats * 495 commits Files Permalink Failed to load latest commit information. Type Name Latest commit message Commit time .github/workflows docs examples siuba .gitignore .travis.yml CODEOWNERS LICENSE MANIFEST.in Makefile README.md docker-compose.yml pytest.ini requirements-dev.txt requirements-test.txt requirements.txt setup.py siuba.Rproj View code README.md siuba scrappy data analysis, with seamless support for pandas and SQL CI Documentation Status Binder [siuba_small] siuba (Xiao Ba ) is a port of dplyr and other R libraries. It supports a tabular data analysis workflow centered on 5 common actions: * select() - keep certain columns of data. * filter() - keep certain rows of data. * mutate() - create or modify an existing column of data. * summarize() - reduce one or more columns down to a single number. * arrange() - reorder the rows of data. These actions can be preceeded by a group_by(), which causes them to be applied individually to grouped rows of data. Moreover, many SQL concepts, such as distinct(), count(), and joins are implemented. Inputs to these functions can be a pandas DataFrame or SQL connection (currently postgres, redshift, or sqlite). For more on the rationale behind tools like dplyr, see this tidyverse paper. For examples of siuba in action, see the siuba documentation. Installation pip install siuba Examples See the siuba docs or this live analysis for a full introduction. Basic use The code below uses the example DataFrame mtcars, to get the average horsepower (hp) per cylinder. from siuba import group_by, summarize, _ from siuba.data import mtcars (mtcars >> group_by(_.cyl) >> summarize(avg_hp = _.hp.mean()) ) Out[1]: cyl avg_hp 0 4 82.636364 1 6 122.285714 2 8 209.214286 There are three key concepts in this example: concept example meaning verb group_by(...) a function that operates on a table, like a DataFrame or SQL table siu _.hp.mean() an expression created with siuba._, that expression represents actions you want to perform pipe mtcars >> a syntax that allows you to chain verbs with group_by(...) the >> operator See introduction to siuba. What is a siu expression (e.g. _.cyl == 4)? A siu expression is a way of specifying what action you want to perform. This allows siuba verbs to decide how to execute the action, depending on whether your data is a local DataFrame or remote table. from siuba import _ _.cyl == 4 Out[2]: #-== +-#-. | +-_ | +-'cyl' +-4 You can also think of siu expressions as a shorthand for a lambda function. from siuba import _ # lambda approach mtcars[lambda _: _.cyl == 4] # siu expression approach mtcars[_.cyl == 4] Out[3]: mpg cyl disp hp drat wt qsec vs am gear carb 2 22.8 4 108.0 93 3.85 2.320 18.61 1 1 4 1 7 24.4 4 146.7 62 3.69 3.190 20.00 1 0 4 2 .. ... ... ... ... ... ... ... .. .. ... ... 27 30.4 4 95.1 113 3.77 1.513 16.90 1 1 5 2 31 21.4 4 121.0 109 4.11 2.780 18.60 1 1 4 2 [11 rows x 11 columns] See siu expression section here. Using with a SQL database A killer feature of siuba is that the same analysis code can be run on a local DataFrame, or a SQL source. In the code below, we set up an example database. # Setup example data ---- from sqlalchemy import create_engine from siuba.data import mtcars # copy pandas DataFrame to sqlite engine = create_engine("sqlite:///:memory:") mtcars.to_sql("mtcars", engine, if_exists = "replace") Next, we use the code from the first example, except now executed a SQL table. # Demo SQL analysis with siuba ---- from siuba import _, group_by, summarize, filter from siuba.sql import LazyTbl # connect with siuba tbl_mtcars = LazyTbl(engine, "mtcars") (tbl_mtcars >> group_by(_.cyl) >> summarize(avg_hp = _.hp.mean()) ) Out[4]: # Source: lazy query # DB Conn: Engine(sqlite:///:memory:) # Preview: cyl avg_hp 0 4 82.636364 1 6 122.285714 2 8 209.214286 # .. may have more rows See querying SQL introduction here. Example notebooks Below are some examples I've kept as I've worked on siuba. For the most up to date explanations, see the siuba docs * siu expressions * dplyr style pandas + select verb case study * sql using dplyr style + simple sql statements + the kitchen sink with postgres * tidytuesday examples + tidytuesday is a weekly R data analysis project. In order to kick the tires on siuba, I've been using it to complete the assignments. More specifically, I've been porting Dave Robinson's tidytuesday analyses to use siuba. Testing Tests are done using pytest. They can be run using the following. # start postgres db docker-compose up pytest siuba About Python library for using dplyr like syntax with pandas and SQL siuba.readthedocs.io Topics python sql dplyr pandas data-analysis Resources Readme License MIT License Releases 18 SQL and fast grouped method improvements Latest Aug 29, 2020 + 17 releases Packages 0 No packages published Used by 20 * @dcamposliz * @CodeForPhilly * @Parth099 * @chriscardillo * @chriscardillo * @lucasmuchaluat * @pwwang * @machow + 12 Contributors 7 * @machow * @breichholf * @kirillseva * @bakera81 * @ismayc * @tmastny Languages * Python 99.7% * Makefile 0.3% * (c) 2021 GitHub, Inc. * Terms * Privacy * Security * Status * Docs * Contact GitHub * Pricing * API * Training * Blog * About You can't perform that action at this time. You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session.