Posit PBC
tidymodels: Adventures in Rewriting a Modeling Pipeline - posit::conf(2023)
updated
Simplify Python web app development with Shiny Express and Posit Team:
Anonymous questions: pos.it/demo-questions
Demo: 11 am ET
Q&A: ~11:30 am ET
Join Winston Chang at Posit on Wednesday, February 28th at 11 am ET to learn more about creating interactive data dashboards and data-driven applications faster than ever.
Live Q&A at ~11:30 am ET will be here: youtube.com/live/zg4LP4lkihM?feature=share
Helpful resources:
🖇️ Introducing Shiny Express: shiny.posit.co/blog/posts/shiny-express
🖇️ Shinylive: shinylive.io/py/examples
🖇️ Component Gallery: shiny.posit.co/py/components
🖇️ VS Code Shiny extension: marketplace.visualstudio.com/items?itemName=Posit.shiny-python
📦 rsconnect-python package: pypi.org/project/rsconnect-python
🖇️ Follow-up links:
Event Survey: https://forms.gle/8AcMAnjfPFSTQZJK6
Posit Team: posit.co/products/enterprise/team
Request evaluation: pos.it/chat-with-us
Posit Team demo resources: pos.it/demo-resources
What will you learn?
If you’re new to Shiny → you’ll get started writing and deploying your first Shiny Express application in Python.
If you already know Shiny → you’ll see how Shiny Express can make your development experience quicker and more efficient.
Happy with the way things are? No need to change what you’re doing! We think of Shiny Express and Shiny Core as complementary, and intend to support both syntaxes indefinitely. Last month’s Workflow with Posit Team Demo featured Shiny in R and used bslib for custom theming, which you can check out here in the recording: youtu.be/O6WLERr5bKU?feature=shared
LIVE Q&A ROOM for ~11:30 am on Feb 28th: youtube.com/live/zg4LP4lkihM?feature=share
There is no need to register; join us here on YouTube at the time above or you can add to your calendar using the link below:
pos.it/team-demo
We host these Workflow Demos on the last Wednesday of every month, so you can use the link above to add the recurring event as well.
Brad Zielke is an Operations Data Science leader focused on unlocking the retail experience.
________________________
► Subscribe to Our Channel Here: bit.ly/2TzgcOu
Follow Us Here:
Website: posit.co
LinkedIn: linkedin.com/company/posit-software
To join future data science hangouts, add to your calendar here: pos.it/dsh (All are welcome! We'd love to see you!)
Thanks for hanging out with us! 💛
Speaker bio: Jamie Warner is an analytics and data science leader with a passion for revolutionizing the way heavily regulated industries understand and adopt data science. She has launched and led data science organizations at multiple Fortune 200 insurers and driven the creation of large-scale pricing and reserving models for insurers. She holds an M.S. in Business Analytics from Bentley University, a B.A. in Mathematics and Economics from Colby College, and earned CPCU and AIDA designations. Jamie is passionate about staying on top of new trends while educating and empowering analytics professionals, so she teaches as an adjunct professor for Northeastern University’s Master’s of Analytics program. She is currently the vice chair of the ethics committee and the immediate past president for the Boston Chapter of the CPCU where she has served on the board of directors since 2015.
________________________
► Subscribe to Our Channel Here: bit.ly/2TzgcOu
Follow Us Here:
Website: posit.co
LinkedIn: linkedin.com/company/posit-software
To join future data science hangouts, add to your calendar here: pos.it/dsh (All are welcome! We'd love to see you!)
Thanks for hanging out with us! 💛
Links mentioned:
🔗 FDA Digital Transformation Symposium: fda.gov/news-events/fda-meetings-conferences-and-workshops/2023-fda-digital-transformation-symposium-12042023
🔗 precisionFDA: precision.fda.gov
🔗 R4DS Online Learning Community: r4ds.io/join
🔗 R Validation Hub Contact: pharmar.org/contact
🔗 Hangout LinkedIn Group: linkedin.com/groups/12610075
Speaker bios:
Kevin Snyder - Associate Director of Nonclinical Informatics at U.S. FDA
Kevin received his Bachelors in biochemistry from the University of Maryland in 2008 and his PhD in neuroscience from the University of Pennsylvania School of Medicine in 2013. He currently serves as the Associate Director of Nonclinical Informatics in the Office of New Drugs in the Center for Drug Evaluation and Research at the US FDA where he manages data science and informatics initiatives to support the pharmacology/toxicology review program. These initiatives include research efforts to develop methods to optimize the regulatory use of standardized electronic CDISC-SEND-formatted toxicology study data as well as internal informatics projects to promote the development of scientifically sound, data-driven regulatory policies. Dr. Snyder also leads an agency-wide Data Science and Software Development working group that is focused on building out the organizational infrastructure necessary to support the work of data scientists across the agency and is an active collaborator with several consortia efforts, e.g. CDISC, PHUSE, and BioCelerate, to improve the implementation and use of the SEND data standard.
Raju (Rama) Rayavarapu - Data Scientist at U.S. FDA
Raju comes to the FDA from the great state of Pennsylvania via South Bend, IN and Memphis, TN where he did his Ph.D. and post-doctoral work. He is the lead of ODAR’s DataForward Initiative which is focused on upskilling FDA staff in data related skills and harnessing the power of the incredible existing FDA data science community to drive data literacy and the joy that comes from working with and understanding data.He loves to spend his time talking FOSS (Python/R/whatever), Natural Language Processing, artificial Intelligence (AI), and all things data science. He also spends any night he can staring into the universe with his two dogs and his telescope.
________________________
► Subscribe to Our Channel Here: bit.ly/2TzgcOu
Follow Us Here:
Website: posit.co
LinkedIn: linkedin.com/company/posit-software
To join future data science hangouts, add to your calendar here: pos.it/dsh (All are welcome! We'd love to see you!)
Thanks for hanging out with us! 💛
Speaker bio: Wes is an entrepreneur and open source software developer focusing on data science tools and analytical computing. He’s currently a Principal Architect at Posit PBC. Previously, Wes co-founded Voltron Data and created or co-created the pandas, Apache Arrow, and Ibis projects. He is a Member of The Apache Software Foundation and has published three editions of Python for Data Analysis.
Links mentioned:
🔗 Python for Data Analysis: wesmckinney.com/book
🔗 Wes McKinney Blog: wesmckinney.com/archives
🔗 Automate the Boring Stuff with Python: automatetheboringstuff.com
🔗 posit::conf(2024) call for talks: posit.co/blog/speak-at-posit-conf-2024
🔗 Quarto: quarto.org
🔗 This is Water: https://fs.blog/david-foster-wallace-this-is-water/
🔗 Questions about Posit Products: https://pos.it/chat-with-us
________________________
► Subscribe to Our Channel Here: bit.ly/2TzgcOu
Follow Us Here:
Website: posit.co
LinkedIn: linkedin.com/company/posit-software
To join future data science hangouts, add to your calendar here: pos.it/dsh (All are welcome! We'd love to see you!)
Thanks for hanging out with us! 💛
Speaker bio: Gerard Sentveld is the Director of Data Analytics Operational Risk Management at Prudential. Gerard has over 25 years of experience as a data expert in a variety of industries from local government, utilities, telecom, pharmaceuticals and most recently the insurance industry. Gerard is currently responsible for managing data science initiatives that quantify operational risk for his stakeholders. In previous roles he hosted data science competitions and connected data scientists from across the enterprise with each other. Gerard earned his MS in Computer Science from the Vrije Universiteit in Amsterdam. Outside of work he is always trying to improve his sound system and find new music to listen to.
________________________
► Subscribe to Our Channel Here: bit.ly/2TzgcOu
Follow Us Here:
Website: posit.co
LinkedIn: linkedin.com/company/posit-software
To join future data science hangouts, add to your calendar here: pos.it/dsh (All are welcome! We'd love to see you!)
Thanks for hanging out with us! 💛
Featured Leader Bio:
Mehran Moghtadai is a Senior Manager at TD Bank leading the AI/ML accelerator and enablement team. Leading the creation and maintenance of numerous internally developed tools. Mehran joined TD Insurance as an actuarial analyst in 2014 working on various modeling projects as well as the implementation of the first optimized pricing algorithm at TD. Mehran holds a Masters in actuarial mathematics and finance from Concordia University.
________________________
► Subscribe to Our Channel Here: bit.ly/2TzgcOu
Follow Us Here:
Website: posit.co
LinkedIn: linkedin.com/company/posit-software
To join future data science hangouts, add to your calendar here: pos.it/dsh (All are welcome! We'd love to see you!)
Thanks for hanging out with us! 💛
Theming Shiny apps with your company brand" session: youtu.be/O6WLERr5bKU
Anonymous questions: pos.it/demo-questions
Demo: 11 am ET
Q&A: ~11:40 am ET
Join Garrett Grolemund at Posit on Wednesday, January 31st at 11 am ET to learn how to theme and brand your own apps.
The session will highlight how to:
✨ Layout an app with the bslib package (modern UI toolkit with no knowledge of CSS required)
✨ Add cards, value boxes, and logos
✨ Customize the theme of the app
✨ Tweak the theme by swapping out primary colors, secondary colors, and more.
✨ Quickly apply the theme to every plot in the app
✨ Work with bootstrap classes
Helpful Resources:
🖇️ Code at : github.com/garrettgman/shiny-styling-demo
📦 bslib package: rstudio.github.io/bslib
📦 bsicons package: github.com/rstudio/bsicons
🖇️ Bootstrap icons: icons.getbootstrap.com
🖇️ Bootstrap CSS classes: bootstrapshuffle.com/classes
📦 thematic package: rstudio.github.io/thematic
📦 gitlink package: github.com/colearendt/gitlink
🖇️ Follow-up links:
Posit Team: posit.co/products/enterprise/team
Request evaluation: pos.it/chat-with-us
Posit Team demo resources: pos.it/demo-resources
LIVE Q&A ROOM for ~11:45 am on January 31st: youtube.com/live/1G8ZM6kbt8c?feature=share
There is no need to register; join us here on YouTube at the time above or you can add to your calendar using the link below:
pos.it/team-demo
We host these Workflow Demos on the last Wednesday of every month, so you can use the link above to add the recurring event as well.
Learn more in our blog post: posit.co/blog/github-copilot-on-posit-cloud
Posit Cloud: https://posit.cloud/
GitHub Copilot: github.com/features/copilot
RStudio User Guide: docs.posit.co/ide/user/ide/guide/tools/copilot.html
Have you ever submitted a report or other product and wished you'd been given just a bit more time to clean up the look of it? Does it make your skin crawl to hear ""the data speaks for itself, don't waste your time making that deliverable 'pretty'""?
As a compliment to the many sessions this week in which you'll hear great methods for *how* to make your work more beautiful, in this talk, we'll walk through some of the scientific research that shows *why* taking the time to make design improvements is critical to communicating your point with data‚ for dashboards, reports, and even simple tables.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Compelling design for apps and reports.
Session Code: TALK-1102
This talk shares practical tips and tangible stories for how intentional approaches to documenting things is helping big distributed teams tackle hard challenges and change organizational culture via NASA Openscapes, NOAA Fisheries Openscapes, & beyond.
I'll share about documenting things, and how intentional approaches to documentation and onboarding are helping big distributed teams tackle hard challenges and change organizational culture. The goal is to provide concrete tips to help you document things effectively & hear stories of how putting a focus on documentation can be help teams be efficient, productive, and less lonely. I'll give a short lightning talk (inspired by Jenny Bryan's Naming Files talk) followed by stories from NASA Openscapes, NOAA Fisheries Openscapes & beyond.
Materials:
- Slides: openscapes.github.io/documenting-things
- Blog post: openscapes.org/blog/2023-09-27-documenting-things-posit-conf
- Website: openscapes.org - links to NASA Openscapes and NOAA Fisheries Openscapes and beyond
- Jenny Bryan's Naming Files talk - github.com/jennybc/how-to-name-files#how-to-name-files
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Getting %$!@ done: productive workflows for data science.
Session Code: TALK-1092
I asked the community if anyone had something to share on this topic, and they delivered! A special thank you to Derek Beaton, Director of Advanced Analytics at St. Michael’s Hospital and Joe Powers, Principal Data Scientist at Intuit who kicked off the group conversation last month by sharing their own experiences.
Tips for communicating the ROI for data science projects:
✨ Engage stakeholders & identify gatekeepers
✨ Align evaluation metrics to stakeholder objectives
✨ Establish clear success criteria
✨ Anticipate where things can go awry
✨ Consistently communicate your central message
✨ Prioritize what’s most impactful
✨ Promote transparency while measuring project performance
✨ Ensure the benefit is unquestionably worth the time investment
Additional ROI Resources Mentioned:
✨ Analytics Power Hour Podcast - Estimating the Effort for Analytics Projects: analyticshour.io/2023/10/31/231-estimating-the-effort-for-analytics-projects
✨ CDO Matters Podcast with Malcom Hawker: open.spotify.com/show/2OJXB8v32CrKbbsv9Uq2vn
✨ How to Measure Anything by Douglas Hubbard: hubbardresearch.com/7-simple-principles-for-measuring-anything
If you want to continue conversations like this, we invite you to also join us at the Data Science Hangout every Thursday from 12-1 ET. Every week we're joined by a different data leader from the community to share their perspectives and experience.
You can add it to your calendar using this link: pos.it/dsh
Featured Leader Bio:
Sean Nguyen is a Senior Staff Data Scientist at S2G Ventures, where he works on the data science and technology team. He focuses on developing data products and identifying actionable insights for the firm and portfolio companies. Before S2G, Sean was a data science fellow at Insight Data Science where he developed a machine learning model to predict the outcome of trademark infringement lawsuits.
Sean received his B.S. from The University of Michigan-Dearborn and his Ph.D. in Cell & Molecular Biology and Environmental Toxicology from Michigan State University. During his time at MSU, he worked as a data scientist in the graduate school and as a commercialization intern in the technology transfer office.
________________________
► Subscribe to Our Channel Here: bit.ly/2TzgcOu
Follow Us Here:
Website: posit.co
LinkedIn: linkedin.com/company/posit-software
Twitter: twitter.com/posit_pbc
To join future data science hangouts, add to your calendar here: pos.it/dsh (All are welcome! We'd love to see you!)
Thanks for hanging out with us! 💛
Speaker Bio:
Matthew McDonald is a Senior Managing Director at Kroll Bond Rating Agency, responsible for managing the Quantitative Modeling team. Matt joined KBRA in 2015 to Build out KBRA’s Model Risk Management framework. Before joining KBRA, Matt held various modeling roles at GE Capital, IBM Global Financing, priceline.com and PriceWaterhouseCoopers. Matt holds Masters degrees from Columbia University and the University of Connecticut, and a BA in Mathematics from Colgate University.
___________________
► Subscribe to Our Channel Here: bit.ly/2TzgcOu
Follow Us Here:
Website: posit.co
LinkedIn: linkedin.com/company/posit-software
Twitter: twitter.com/posit_pbc
To join future data science hangouts, add to your calendar here: pos.it/dsh (All are welcome! We'd love to see you!)
Thanks for hanging out with us! 💛
By popular demand, our upcoming monthly workflow with Ryan Johnson on December 27th is dedicated to enhancing teamwork. It's recorded, so no worries if you're out! Or perhaps you’ll add us to the mix of holiday movies and watch from the couch!
This Month's Focus: All Things Collaborative Working
🔄 Version control
💻 Git-backed deployment
👥 Project sharing
Date & Time: Wednesday, December 27th at 11 am ET.
📦 Packages mentioned:
Shiny: shiny.posit.co
bslib: rstudio.github.io/bslib
🖇️ Follow-up links:
Posit Team: https://posit.co/products/enterprise/...
Talk to us directly: pos.it/chat-with-us
Posit Team demo resources: pos.it/demo-resources
There is no need to register; join us here on YouTube at the time above or you can add to your calendar using the link below:
pos.it/team-demo
We host these Workflow Demos on the last Wednesday of every month, so you can use the link above to add the recurring event as well.
We will use this thread on the Posit Community Forum for follow-up Q&A from this month's session: community.rstudio.com/t/event-on-12-27-collaborative-workflows-w-posit-team-version-control-git-backed-deployment-project-sharing/179181 (shortlink: pos.it/workflow-dec-23)
Happy holidays! Cheers to 2024!
Speaker bio:
Laura Gast is an epidemiologist specialized in simplifying complex data for high-level decision-making, constantly trying to 'solve' the idea in information communication to "Make things as simple as possible, but no simpler." Laura has a strong background in infectious and vector-borne diseases, gained while shaping malaria elimination strategies in Southern Africa and contributing to public health projects across at least ten countries. She started out in the AI/ML space in the early 2000s with her doctoral research which employed NASA satellite imagery and machine learning to analyze land-use changes in relation to mosquito-borne diseases in Peru.
Currently, she serves at a nonprofit focused on the well-being of enlisted U.S. military members and their families, tackling data challenges in development and impact measurement. Outside of her day job, Laura relishes pub trivia, live comedy or music performances, and long runs on the National Mall.
______________
► Subscribe to Our Channel Here: bit.ly/2TzgcOu
Follow Us Here:
Website: posit.co
LinkedIn: linkedin.com/company/posit-software
Twitter: twitter.com/posit_pbc
To join future data science hangouts, add to your calendar here: pos.it/dsh (All are welcome! We'd love to see you!)
Thanks for hanging out with us! 💛
Speaker bio:
Stephanie Lussier is a Manager of Biostatistics at Moderna. She’s a member of the Moderna Specialty Data Analytics team, which focuses on building a seamless analytic platform and capabilities to enable quantitative decision-making in a timely fashion across clinical programs. Her professional interests include designing novel data visualizations, building tools that help clinical trial statisticians, and promoting the use of open-source software.
______
► Subscribe to Our Channel Here: bit.ly/2TzgcOu
Follow Us Here:
Website: posit.co
LinkedIn: linkedin.com/company/posit-software
To join future data science hangouts, add to your calendar here: pos.it/dsh (All are welcome! We'd love to see you!)
Thanks for hanging out with us! 💛
This is the story of how a Royal Statistical Society writer discovered Quarto, learned how to code (a bit), and built realworlddatascience.net, an online publication for the data science community.
In March 2022, I was tasked by the Royal Statistical Society with creating a new online publication: a data science website for data science professionals. I've been a print journalist for 20 years and have worked on websites in that time, but my coding ability began and ended with wrapping href tags around text and images. That is until I discovered Quarto. In this talk, I describe how I explored, learned, and fell in love with the Quarto publishing system, how I used it to build a website -- Real World Data Science (realworlddatascience.net) -- and how the open source community mindset helped shape my thinking about what a new publication could and should be!
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Quarto (1).
Session Code: TALK-1071
In this talk, I'm sharing my personal journey as a data scientist and the key lessons learned along the way. I'll emphasize the importance of finding a positive community of like minded-allies, persevering through setbacks as success is not linear, and exploring by embracing the broad nature of the data science field. By sharing my experiences and acknowledging the challenges I've faced attendees will gain a fresh perspective on what it takes to succeed in a data science career and inspire them to pursue their passions in the field.
Overall, this talk aims to provide a glimpse into the reality of a data science career. Attendees will take away a sense of motivation and empowerment to find their own unique path to success.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Lightning talks.
Session Code: TALK-1169
Data scientists are creating incredibly useful data products at an accelerating rate. These products are consumed by others who expect them to be accurate reliable and timely, often promises unfulfilled. In this talk, we will explore how to use common CI/CD pipeline tools already within reach of attendees to automatically test and deploy their apps, APIs, and reports.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Lightning talks.
Session Code: TALK-1166
I will share what our team has learned from successfully integrating R in all areas of our company's operations. InsightRX is a precision medicine company whose goal is to ensure that each patient receives the right drug at the optimal dose. At InsightRX, R is a first-class language that's used for purposes ranging from customer-facing products to internal data infrastructure, new product prototypes, and regulatory reporting. Using R in this way has given us the opportunity to forge fruitful collaborations with other teams in which we can both learn and teach.
Join me as I share how the skills of data science and engineering can complement each other to create better products and greater impact.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: R Not Only In Production.
Session Code: KEY-1108
starts_with(language): Translating select helpers to dbt. Translating syntax between languages transports concepts across communities. We see a case study of adapting a column-naming workflow from dplyr to dbt's data engineering toolkit.
dplyr's select helpers exemplify how the tidyverse uses opinionated design to push users into the pit of success. The ability to efficiently operate on names incentivizes good naming patterns and creates efficiency in data wrangling and validation.
However, in a polyglot world, users may find they must leave the pit when comparable syntactic sugar is not accessible in other languages like Python and SQL.
In this talk, I will explain how dplyr's select helpers inspired my approach to 'column name contracts,' how good naming systems can help supercharge data management with packages like {dplyr} and {pointblank}, and my experience building the {dbtplyr} to port this functionality to dbt for building complex SQL-based data pipelines.
Materials:
- github.com/emilyriederer/dbtplyr
- emilyriederer.com
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Databases for data science with duckdb and dbt.
Session Code: TALK-1098
The development of software can be costly and time-consuming. If the end users are not involved in the process from the start the tool we built may not meet their needs. In this presentation, I will discuss how prototyping in Shiny can help you build the right tool and save you from spending millions of dollars on a tool no one will use. I will explore the advantages of using Shiny for prototyping, particularly its ability to rapidly build interactive applications. I will also discuss how to design effective prototypes, including techniques for gathering user feedback and using that feedback to refine your tool. I will emphasize the importance of presenting real-life data, particularly when building data-driven tools.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Shiny user interfaces.
Session Code: TALK-1125
Ten years ago, the first set of git commits was submitted to a new R software package repository "dataRetrieval" with the goal to provide an easy way for R users to retrieve U.S Geological Survey (USGS) water data. At that time, the perception within the USGS was the use of R was exclusive to an elite group of "very serious scientists." Fast forward, we now find many newer USGS hires having a solid grasp of the language from the start along with the use of R in a wide variety of applications.
In this talk, I'll discuss my experiences maintaining the dataRetrieval package, how it's shaped my career, impacted USGS R usage, and why data providers should consider sponsoring their own R packages wrapping their data API services.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Lightning talks.
Session Code: TALK-1171
Onboarding new hires can be a challenging process, but taking a problem-focused approach can make it more meaningful and rewarding. In this talk, I will share how I discovered the ultimate new hire hack by creating two small packages that gave me the confidence I needed when I started at BMS. Through building these packages, I not only learned R things like using bslib and making font files available for published dashboards, but also gained a deep understanding of my company's internal systems and workflows, and connected with my team via lots of questions. The resulting packages are still heavily used today.
Join me to discover how small packages can have a broad impact and what hiring managers can do to help.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Developing your skillset; building your career.
Session Code: TALK-1112
🚀 Elevate your Quarto projects to new heights with these practical tips and tricks!💡
"Wiki", "User Guide", "Handbook" -- whatever you call yours, we converted ours to Quarto!
A year ago, my team's documentation, which had been created using Microsoft Word, was large and lacked version control. Scrolling through the document was slow, and, due to confidentiality reasons, only one person could edit it at a time, which was a significant challenge for our team of multiple developers. After realizing we needed a more flexible solution, we successfully converted our documentation to Quarto.
In this talk, I'll discuss our journey converting to Quarto, the challenges we faced along the way, and tips and tricks for anyone else who might be looking to adopt Quarto too.
Slides: https://melissavanbussel.quarto.pub/posit-conf-2023;
Code for slides: github.com/melissavanbussel/posit-conf-2023;
My YouTube: youtube.com/c/ggnot2;
My website: melissavanbussel.com/;
My Twitter: twitter.com/melvanbussel;
My LinkedIn: linkedin.com/in/melissavanbussel
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Quarto (2).
Session Code: TALK-1140
The pharmaceutical industry is undergoing rapid change, driven by a desire from both industry and regulatory agencies to move to more interactive visualizations and web applications to review data and make decisions. These changes would have been unthinkable 30 years ago when I started working at Pfizer.
In this talk, I'll consider the drivers for these changes, how open-source tools can help achieve this, and why collaboration across the industry is vital to achieving this goal. I'll contrast this with my experience of 30 years working in the pharma industry - when the R language had only just been released, when the internet was new, and when submissions to agencies were printed out, loaded onto trucks, and shipped to their doors.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Pharma.
Session Code: TALK-1067
In this talk we share how good programming practices inspire the way we manage the R-Ladies community in order to make it sustainable.
R-Ladies' first ten years were about growing the community: from being just one chapter in 2012 to becoming a global organization in 2016, and now fostering more than 230 chapters worldwide. But how can we face the challenges of growing an organization based solely on volunteer work?
In this talk, we discuss how good programming practices –such as modularity, refactoring, and testing– inspire the way we see the sustainable management of an ever-growing community. To that end, we will present our most recent efforts at creating and documenting workflows, distributing the workload, and automating tasks that allow volunteers to focus their time where it is most needed.
After watching this talk, you will get some ideas on how to support volunteers in your own community or project, and on how to use GitHub Actions to automate workflows and tasks.
Learn more and join at: rladies.org
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: It takes a village: building and sustaining communities.
Session Code: TALK-1128
(Due to unforeseen circumstances, Hadley Wickham presented this talk "slide karaoke" style, from materials prepared by Jenny Bryan.)
In R, the fundamental unit of shareable code is the package. As of March 2023, there were over 19,000 packages available on CRAN. Hadley Wickham and I recently updated the R Packages book for a second edition, which brought home just how much the package development landscape has changed in recent years (for the better!).
In this talk, I highlight recent-ish developments that I think have a great payoff for package maintainers. I'll talk about the impact of new services like GitHub Actions, new tools like pkgdown, and emerging shared practices, such as principles that are helpful when testing a package.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Package development.
Session Code: TALK-1132
Using Small-Multiples (faceted graphs) is an effective way to compare patterns across many dimensions. In this talk, I'll walk you through some ways to lay out your individual facets according to the underlying data. For example, maybe each facet represents a city or point on a 2D plane - we'll explore ways to organize facets in a grid that mimics the data itself - unlocking your ability to explore patterns in 4+ dimensions. Other solutions to this problem rely on manually-curated lists that map common layouts to a grid, but in this talk, we'll explore solutions that work on EVERYTHING. I'll show you how to incorporate this technique into your viz and how I built the libraries since there are some interesting data science concepts at play.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Lightning talks.
Session Code: TALK-1174
A 5-minute talk to discuss how I've used Quarto and Bootstrap variables to quickly make Shiny's new website look as it should. The Quarto user I have in mind works at an organization with specific brand guidelines to follow. I‚ will discuss how to set up your theme, show some key Quarto settings, and how Bootstrap‚ Sass variables are your best friend.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Lightning talks.
Session Code: TALK-1170
Invasive species are a huge threat to lake ecosystems in Minnesota. With over 10,000 water bodies across the state, having up-to-date data and decision support is critical. Researchers at the University of Minnesota have created four complex R and Python models to support lake managers, all pulled together and presented with the most recent infestation data available.
Come along with us to see how we connected these models in the AIS Explorer, a decision support application built in Shiny to help prioritize risks and placing watercraft inspectors, using tools like OCPU and cloud toolings like Lambda, EventBridge and AWS S3.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: R or Python? Why not both!.
Session Code: TALK-1118
Every data analysis in Python starts with a big fork in the road: which DataFrame library should I use?
The DataFrame Decision locks you into different methods, with subtly different behavior::
- different table methods (e.g. polars `.with_columns()` vs pandas `.assign()`)
- different column methods (e.g. polars `.map_dict()` vs pandas `.map()`)
In this talk, I'll discuss how siuba (a dplyr port to python) combines with duckdb (a crazy powerful sql engine) to provide a unified, dplyr-like interface for analyzing a wide range of data sources‚ whether pandas and polars DataFrames, parquet files in a cloud bucket, or pins on Posit Connect.
Finally, I'll discuss recent experiments to more tightly integrate siuba and duckdb.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Databases for data science with duckdb and dbt.
Session Code: TALK-1101
One in five low-income renter households in the US experienced falling behind on rent or being threatened with eviction in 2021. Yet most are unrepresented when facing eviction in court. The complex and fast-paced legal system obscures access to timely information, leaving tenants without assistance.
In this talk, I discuss the Civil Court Data Initiative's use of R alongside AWS Cloud and SQL to analyze disaggregate eviction records. I focus on the integration of RMarkdown with Amazon Athena and EC2 to create weekly eviction reports across 20 states for legal aid groups working to assist tenants. The upshot: accessible eviction data to help legal aid providers better address local legal needs.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: End-to-end data science with real-world impact.
Session Code: TALK-1146
In the past year, people have come to realize that AI can revolutionize the way we work. This talk focuses on using AI tools with Shiny for Python, demonstrating how AI can accelerate Shiny application development and enhance its capabilities. We'll also explore Shiny's unique ability to interface with AI models, offering possibilities beyond Python web frameworks like Streamlit and Dash. Learn how Shiny and AI together can empower you to do more, and do it faster.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: I can't believe it's not magic: new tools for data science.
Session Code: TALK-1153
The R programming language offers the versatility to perform statistical analyses, create publication-ready plots, and render high-quality reports and presentations. Despite having this environment of indispensable tools, it can be daunting for a beginner-level programmer to get started. Luckily, the Posit community is one of a kind and values inclusivity, collaboration, and empathy. By putting a face to the R packages we use on a daily basis, we hope to make every programmer feel included and capable. We want to inspire attendees to create their own projects or packages, connect with others inside and outside of their field of expertise, and challenge themselves to learn something new, knowing the community is right there to support them.
Materials: http://www.sarmapar.com/people_of_posit
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Lightning talks.
Session Code: TALK-1165
Spark Connect, and Databricks Connect, enable the ability to interact with Spark stand-alone clusters remotely. This improves our ability to perform Data Science at-scale. We will share the work in `sparklyr`, and other products, that will make it easier for R users to take advantage this new framework.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Tidy up your models.
Session Code: TALK-1084
Abstractions rule everything around us. JD Long talks about abstractions from the board room to the silicon.
Over 20 years ago Joel Spolsky famously wrote, "All non-trivial abstractions, to some degree, are leaky." Unsurprisingly this has not changed. However, we have introduced more and more layers of abstraction into our workflows: Virtual Machines, AWS services, WASM, Docker, R, Python, data frames, and on and on. But then on top of the computational abstractions we have people abstractions: managers, colleagues, executives, stakeholders, etc.
JD's presentation will be a wild romp through the mental models of abstractions and discuss how we, as technical analytical types, can gain skill in traversing abstractions and dealing with leaks.
Materials: github.com/CerebralMastication/Presentations/tree/master/2023_posit-conf
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: It's abstractions all the way down ....
Session Code: KEY-1161
At Pfizer, we have over 1500 users with R installed on their machines, along with an R community on MS Teams comprising over a thousand colleagues globally. How can we effectively engage with Pfizer R users and celebrate the successes of this community, while sharing best practices? Additionally, how do we avoid isolated groups duplicating efforts to solve R-related problems across different parts of the organization?
To address these challenges, we established the Pfizer R Center of Excellence (CoE) in early 2022. We focus our efforts on bringing together a rapidly growing community of colleagues, providing technical expertise, and offering best-practice guidance. A well-established, maintained and engaged R community promotes an inclusive and supportive learning environment that drives innovation within organizations. Our aim is to help colleagues thrive in their R journey, regardless of their expertise level.
During my talk, I will share the techniques we used to build a supportive R community, the tools employed to increase community engagement, and the successes and challenges encountered in building an engaging community of R users.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Pharma.
Session Code: TALK-1066
Are you considering or curious about developing code-based tools for scientists? Whether you are an experienced developer or a fellow Posit Academy graduate who might be stepping into this role for the first time, the aim of my story is to inspire you and help you navigate this process. While developing custom R functions, packages, and Shiny apps for diverse analytical capabilities and users in R&D, I learned why it's important to collect certain information at the start before writing any tidying, analysis, visualization, and web application code.
In this talk, I will share the essential technical questions that help me define and plan for success.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: Lightning talks.
Session Code: TALK-1168
Many data science meetup organizers struggle with burnout. It can be daunting to plan a meetup schedule, especially with the added burden of work and life.
In this talk, I want to highlight some strategies for keeping your data science meetup sustainable. Specifically, I want to highlight the role of self-care in growing and sustaining your group, as well as low-key activities like a data scavenger hunt, watching videos together, styling plots together, and sharing useful tidyverse functions.
By making it easy for your members to contribute and empowering them, it takes a lot of the burden off you as an organizer. You don't need to reinvent the wheel for meetups or have famous guests for each one. Let's start the conversation and make your meetup last.
Presented at Posit Conference, between Sept 19-20 2023,
Learn more at posit.co/conference.
--------------------------
Talk Track: It takes a village: building and sustaining communities.
Session Code: TALK-1129


