PyData
NVIDIAs Andy Terrel on GPUs, Open Source Science & Why Code Has Always Been a Tax | GM5 S2 Finale
updated
Talk given by Derek Whitley
Why is uncertainty-awareness important? Well, would you, as a Data Scientist, accept to be evaluated on reducing mean absolute error of some model by 50%? Probably not - but would you feel comfortable explaining why? This talk will establish why such generic and ad hoc goal setting is not meaningful, why model judgement is harder than expected, and how it can still be done reliably and without too many technicalities. We will exemplify uncertainty-aware model rating using the M5 competition data (Walmart sales numbers). Some immediate interpretations (“model A is clearly better than model B”) will turn out to be flawed upon closer inspection, and we’ll see how to correct them, using standard python libraries (numpy, scipy, pandas).
Takeaways:
- We are typically too self-confident in our skills when it comes to judging models. Instead of jumping to immediate conclusions (“that 80%-error model is bad!”), we should take a step back, build a “reasonable best case”, and benchmark the candidate against that.
- Accepting and dealing with uncertainty is strength, neglecting it is madness. Uncertainty-aware model rating allows us to make reliable statements about “how good the model really is”, without “it depends” and “buts”.
- Requirements from statisticians and business stakeholders can be reconciled by taking both of them seriously. Standard python tooling suffices to improve our modeling of what we know that we don't know.
Talk given by Malte Tichy
Talk given by Kaushik Bokka
In this talk, we'll cover the benefits of a data-centric AI approach – spoilers: increased performance is one of them! –, and cover practical tips on how you can integrate data-centric principles in your daily work.
Andrew Ng once said "Data is the food for AI" when talking about data-centric AI. If that's the case, this talk will provide you with the recipes.
Talk given by Marysia Winkels
Talk given by Samuel Oranyeli
In this talk we will dive into Polars, a new dataframe library backed by Arrow and Rust that offers an expressive API for dataframe manipulation with excellent performance.
If you are a seasoned pandas user willing to explore alternatives, or a beginner user wondering what all the fuzz about these new dataframe libraries is, this talk is for you!
Talk given by Juan Luis Cano Rodríguez
Talk given by by Shashank Shekhar
explain a machine learning model by giving an alternate class prediction
of a data point with some minimal changes in its features.
In this talk, we describe a counterfactual (CF)
generation method based on particle swarm optimization (PSO) and how we can have greater control over the proximity and sparsity properties
over the generated CFs.
Talk by Niranjan G S & Shashank Shekhar
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
For the talk info and speaker bio, please see meetup.com/pydatachi/events/288921496
“How Geometry Helps in Data Analysis and ML Problems.”
Many ML models work with geometrical representations of data, from simple linear regression to convoluted NN algorithms sending images to vectors in unspeakably high-dimensional feature space. That makes geometry and, in particular, differential geometry a powerful tool to tackle ML problems.
Denis Fedoseev gives a brief outline of the geometrical concepts and methods which can be used in this context, including:
-the general notion of manifold (which will help to understand the Manifold hypothesis in data science)
-tangent space
-Riemannian metric
-Minkowski dimension
===
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Near-Duplicate Ad Detection in Online Classified Ad Services
Near-Duplicate Ads are harmful to online classified ad services in many ways.
Demand-side users face low-quality listings, which increases the time and effort required to find desired ads. Also, normal supply-side users get fewer views and make less profit out of their ads. Finally, Duplicate ads may cause a direct decrease in the business’s revenue (By skipping payments such as ad boosting)
In this talk, first, we will discuss the problem, definition, metrics, and training data generation. Then, we will talk about modeling, feature engineering, and how step-by-step metrics were improved.
For texts, we have tried different approaches such as MinHash, CountVectorizer, Bi-LSTM, and transformers. For images, we have tried different approaches, such as Perception-Hash and CNNs.
Also, we will discuss how to apply these approaches to find global duplicate ads and confront spammers.
Presentation Slides: http://pydatayerevan.aua.am/files/2022/09/Diyar-Mohammadi.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
EENLP: Cross-lingual Eastern European NLP Index
In our project we present a wide index of existing Eastern European language datasets (90+) and models (60+). Furthermore, to support the evaluation of commonsense reasoning tasks, we compile and publish cross-lingual datasets for five such tasks and provide evaluation results for several existing multilingual models.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Andrey-Manoshin.pdf
--
Hayk Aprikyan Presents:
What can your Telegram tell about you? (Answer: Everything)
How much has your vocabulary changed over the last year? Who shares the funniest memes with you? And does she find you interesting to chat with? ̶N̶o̶p̶e̶.̶
If you're a Telegram guy, Neplo is that painstakingly data-driven guy who's got answers to these (and hundreds of other) questions based on your Telegram chat histories.
Still skeptical? Come and see. (John 1:39)
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Hayk-Aprikyan.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Large Scale Representation Learning In-the-wild
A significant amount of progress is being made today in the field of representation learning. It has been demonstrated that unsupervised techniques can perform as well as, if not better than, fully supervised ones on benchmarks such as image classification, while also demonstrating improvements in label efficiency by multiple orders of magnitude. In this sense, representation learning is now addressing some of the major challenges in deep learning today. It is imperative, however, to understand systematically the nature of the learnt representations and how they relate to the learning objectives.
In this talk, we will present a comprehensive overview of representation learning from the beginning to the modern models, contextualize these methods, and discuss the pros and cons of current evaluation methods. New era of deep learning methods that could understand simultaneously the different variety of data will introduce by its evolutions.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Representation-Learning-Presentation.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Being part of statistical learning apparatus, and having a strong mathematical background AB testing remains one of the aspects in the field that continue to be violated and misinterpreted. A big part of violations covers the wrong experiment setup, which I'll try to cover in practice taking into consideration the business setup: whether it's a B2B platform or B2C. It would be nice if the audience had a hands-on experience with A/B testing, if not - I'm still going to cover it on a high level. The main takeaway for the audience will be to understand the pitfalls that relay under experiment setup, where a single disregarded use case can violate the whole experiment outcome.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Sona-Hambaryan.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Target-Based Sentiment Analysis with T5
The classic sentiment analysis analyzes texts, images, emojis, etc to know what other people think of a product, service, company, or event. While sentiment analysis can be considered one of the accomplished tasks of Natural Language Processing tasks, more fine-grained types of it like Target Based Sentiment Analysis(TSA) or Aspect-based sentiment analysis(ABSA) are not quite the same. In TSA we want to see the sentiment of a given text towards a particular entity(in my case person or organization). This task is one of the non-solved ones. With the T5 question answering transformer model it was possible to solve the task with results 20% higher than the current leaderboards.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Liana-Minasyan.pptx
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Bachelor theses in Deep Learning: Submitted to an Armenian University
Bachelor theses written in the area of Deep Learning based object detection will be presented. The main focus is on the detection of vehicles captured from the top, e.g. parking lots, satellites: I will present the challenges we encountered and solved in the scope of Bachelor thesis. The goal of this talk is not only to present the results of young undergraduate students but also to encourage new ones to get involved in the sphere.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Sergey-Hayrapetyan.pptx
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
NetworkX - your Unexpected Assistant for Clustering Analysis
Clustering analysis is a common task in data science but it can sometimes get tedious. In this talk, I will present how functionality from the package NetworkX can assist us in analyzing and presenting the results of clustering analysis. This talk assumes no previous knowledge, a brief reminder of graph basics will be given and networkx will be shortly presented.
Join in if you:
- want to hear about a new suggested usage of a known data structure,
- like to get things done more efficiently when you cluster,
- never heard of networkx but would like to,
- all of the above.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Meirav-Ben-Izhak.pptx
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Large Scale Field Delineation
The talk is about methods of doing large-scale field delineation from aerial imagery. Given the increasing importance of global food supplies, AI in agriculture has become integral for later development in the field. Given that modern agriculture is field level, a delineation of fields is required to be able to use such methods.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Hrach-Asatryan.pptx
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Scaling Semi-Supervised Production-Grade ASR on 200 Languages
Self-Supervised pretraining has been wildly successful lately, covering almost every domain: Speech, NLP, Vision. Networks, such as: Wav2Vec2, Hubert, JUST, and alikes have enabled rapid development of Speech-related products. In this talk, we're going to go through the end-to-end research and engineering process of production-grade self-supervised ASR in the multilingual setting. Covered topics include: Compute, Data, Scalability, Engineering for Pretraining, and Downstream Tuning.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Luka-Chkhetiani.pptx
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Building a Lakehouse data platform using Delta Lake, PySpark, and Trino
In this talk I would like to present the concept of Lakehouse, which is a novel architecture to resolve problems and combine capabilities of the classical Data Warehouse and Data Lake. I will talk about the Delta Lake table format that resides in the core of Lakehouse. I will demonstrate how Delta Lake integrates with Apache Spark, to build data ingestion pipelines. I will also show how Delta Lake integrates with Apache Trino, to provide a fast SQL-based serving layer. As a result, I will bring all these components together to describe how they enable a modern big data platform.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Viacheslav-Inozemtsev.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Building a Streaming (E-health) Data Pipeline: When and How?
In this talk, we discuss streaming and the real-time data stack as a solution for analyzing massive, unbounded data sets that are increasingly common in many modern businesses in different fields and their need for more timely and accurate answers.
A streaming data pipeline flows data continuously from source to destination as it is generated, making it being processed along the way so they are used when the analytics, application, or business process requires an updating data flow for an on-time analysis. This analysis can be descriptive like a data dashboard, diagnostic like monitoring logs, predictive like an online fraud detection system, or prescriptive like data process in self-driving cars.
During the talk, examples of e-health data pipelines plus our experience in setting up a data streaming stack help to explicitly of the subject.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Zohreh-Jafari.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Classical Texture Synthesis and Beyond
Given the structural definition of a texture as a special variation in diverse layers of pixels demonstrating reiterating patterns combined with varied randomness in quantity, the purpose of texture synthesis is to generate an expanded vision of the input texture that perceptually resembles the input. The goal of the talk is to provide an overview of classical and neural texture synthesis algorithms. First, two classical non-parametric methods namely Texture Optimization for Example-based Synthesis and Image Quilting for Texture Synthesis and Transfer are covered. Second, two neural methods of texture synthesis are discussed: Texture Synthesis using CNNs and Non-Stationary Texture Synthesis by Adversarial Expansion. Third, the advantages and disadvantages of these four methods are demonstrated. Results of the indicated four approaches and a visual comparison are provided.
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
BERT Model for Real World Healthcare Data
Early indication and detection of diseases, can provide patients with the chance of early intervention, better disease management, and efficient allocation of healthcare resources. The latest developments in machine learning provide a great opportunity to address this unmet need. In this lecture, we introduce modified BERT: A deep neural sequence transduction model designed for electronic health records (EHR). We will consider the application of this methodology to the task of classifying patients into cohorts reflecting different disease patterns.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Artem-Terentyuk.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
NVIDIA NeMo: Toolkit for Conversational AI
Conversational AI is a technology that allows a “machine” to speak to a person in a natural language. NVIDIA NeMo is an open-source conversational AI toolkit built for researchers working on automatic speech recognition (ASR), natural language processing (NLP), and text-to-speech synthesis (TTS). The primary objective of NeMo is to help researchers from industry and academia to develop new models for automatic speech recognition, text-to-speech, natural language processing, and neural machine translation. Nemo also has a large number of step-by-step tutorials and pre-trained models.
The outline of the talk goes as follows:
1. NeMo overview.
2. Where to start: tutorials on ASR, TTS, and NLP.
3. NeMo ASR overview.
4. NeMo TTS overview.
5. NeMo NLP overview.
6. From research to production: deploying NeMo models.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Alex-Laptev.pptx
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
AI-Powered Solutions for Cybersecurity
Cyberattacks are continuously growing in volume and entanglement. They target organizations' systems, networks, and private data, causing financial loss, customer loss, and data leakage. As technology improves nowadays, Artificial Intelligence (AI) based solutions help boost Cybersecurity. This talk will discover how AI-powered algorithms are used to stay ahead of Cyberattacks such as Phishing, Lookalike domains, or Name Spoofing.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Elina-Israyelyan.pptx.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
The Explainability Problem: Towards Understanding Artificial Intelligence
This talk discusses Explainable AI using examples of interest for both machine learning practitioners and non-technical audiences. This talk is not very technical; it does not focus on how to apply an existing method to their model. Rather, the talk discusses the problem of Explainability_ as a whole, namely: what is the Explainability Problem and why it must be solved, how recent academic literature addresses the problem, and how the problem will evolve with new legislation.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/The-Explainability-Problem_-Towards-Understanding-Artificial-Intelligence.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Using Few Shot Object Detection for Utility Pole Detection from Google Street View images
Traditional methods of detecting and mapping utility poles are manual, time-consuming, and costly processes. Current solutions focus on detection of T-shaped (cross-arm-shaped poles) and the lack of labeled data makes it difficult to generalize the process of other types of poles. This work aims to use Few Shot Object Detection techniques to overcome the unavailability of the data and to create a general pole detection model with few labeled images.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Mark-Hamazaspyan.pptx
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
How to Start Critical Thinking in Data Science
The aim of the presentation is to address issues concerning bias in data, misleading statistics, issues in testing, and other matters that are prevalent in the field of data science.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Arpi-Sahakyan.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Sequential Attention-Based Neural Machine Translation
The sequence-to-sequence models can be augmented using an attention mechanism. This algorithm will help your model understand where it should focus its attention given a sequence of inputs. This tutorial will introduce you to sequential and attention models by utilizing the neural machine translation (NMT) model implementation from scratch. Several sequence-to-sequence architectures will be presented under the attention models, including basic models, and intuitions under the attention model. We will then implement the model step by step together and see whether we can figure out the kind of model to translate (or transform) sequences of data such as texts and speech.
Github Page: github.com/hkhojasteh
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Empirical Determinacy of Posterior Location and Scale in Bayesian Hierarchical Models
The parameters in a statistical model are not always identified by the data. In Bayesian analysis, this problem remains unnoticed because of prior assumptions. It is crucial to find out whether the data determine the posterior parameters. In particular, it is important to learn to what extent the spread and the location of the marginal posterior distribution of the parameters are determined by the data.
The R package ed4bhm allows to investigate the empirical determinacy of marginal posterior parameters, their spread, and location. During this talk, I will showcase the functionality of the package ed4bhm with an application of Bayesian logistic regression to Bacterial resistance data.
In this talk, you will learn about:
1. The typical problems in a Bayesian Hierarchical Model (BHM)
2. The theory behind the empirical determinacy of posterior parameters in BHMs
3. How to fit basic BHM in R
4. How to apply the R package ed4bhm and how to interpret the results.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Sona-Hunanyan.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Grover’s Quantum Search for Data Science and Why should we Care
Among the most prominent achievements of the quantum computing field is an algorithm known as Grover’s quantum search. This talk focuses on Grover’s algorithm and its applications to machine learning routines. Prior knowledge required is a basic understanding of linear algebra and computer science, and familiarity with the concepts of machine learning.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Tigran-Sedrakyan.pptx
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Explainable AI as a Conventional Data Analysis Tool
The recent surge of interest in Machine Learning (ML) and Artificial Intelligence (AI) has spurred a wide array of models designed to make decisions in a variety of domains, including healthcare [1, 2, 3], financial systems [4, 5, 6, 7], and criminal justice [8, 9, 10], just to name a few. When evaluating alternative models, it may seem natural to prefer those that are more accurate. However, the obsession with accuracy has led to unintended consequences, as developers often strove to achieve greater accuracy at the expense of interpretability by making their models increasingly complicated and harder to understand [11]. This lack of interpretability becomes a serious concern when the model is entrusted with the power to make critical decisions that affect people’s well-being. These concerns have been manifested by the European Union’s recent General Data Protection Regulation, which guarantees a right to explanation, i.e., a right to understand the rationale behind an algorithmic decision that affects individuals negatively [12]. To address these issues, a number of techniques have been proposed to make the decision-making process of AI more understandable to humans. These “Explainable AI ” techniques (commonly abbreviated as XAI) are the primary focus of this talk.
The talk will be divided into three sections, during which the audience will learn:
(i) the differences between existing XAI techniques,
(ii) the practical implementation of some well-known XAI techniques, and (iii) possible uses of XAI as a conventional data analysis tool.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Maria-Sahakyan.pptx
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Moving Inference to Triton Servers
The talk will introduce the audience to Triton Inference Server, the requirements for migrating from regular AWS instances, advantages and benchmarks of our production.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Marine-Palyan.pptx
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Building your own Multiskill AI Assistant with DeepPavlov
Did you ever dream of having your own AI assistant? Did you find yourself limited by Amazon Alexa or Google Assistant? Did you want to build yours for yourself or your company?
In this talk, you will learn how to build your own multiskill AI Assistant using the modern NLP techniques from DeepPavlov.ai. DeepPavlov is a well-known Conversational AI lab that has participated twice in the Amazon Alexa Prize Challenge, organized Conversational AI Challenges at NeurIPS, and built its very own open-source Conversational AI Stack.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Daniel-Kornev.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Semantic Multimodal Multilingual Similarity Engine
Since multimodality became popular, lots of engineers are trying to make a domain-universal search. The search engines that will find in images by textual query, HTML file by piece of audio, and so on. So here is our (Unum) approach with a bias toward GPU accelerating inference (for underlying models) and a passion to distribute everything.
During the presentation, we will discuss the following questions:
-How to build fast and precise Semantic Textual Similarity engine.
-Multilingual sentence encoders.
-What is CLIP and how can we find an image with a textual query.
-Building indices: Approximate Nearest Neighbor Algorithms and toolkits.
-What is the future of the semantic cross-domain search.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Vladimir-Orshulevich.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Eating humble Py: From toy problem to real-world solution in predicting Customer Lifetime Value
This is the story of my team’s journey from play-problem to real-world solution. Learn, as we learned, what is Customer Lifetime Value and why does everyone in retail suddenly want to predict it? Take a tour of common approaches to solving this problem, from machine learning to good old-fashioned spreadsheets. Feel all the practical pains our clients inflicted on us, and discover why CLV prediction is not as easy as Towards Data Science makes it out to be.
Whether you’re an analytics enthusiast, a novice data scientist or an experienced practitioner, and whether you work in retail or not, there’s something in this talk for you: a little bit of machine learning theory, a peek into a new domain of application you may not be familiar with, or the chance to just cringe in sympathy at problems you know only too well.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Katherine-Munro-.pptx
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Active Learning for 3D Mesh Semantic Segmentation
The talk is about applications of active learning methods, mainly Monte-Carlo Dropout on 3D mesh/pointcloud semantic segmentation task. The topic is particularly interesting for practical applications of Deep Learning models on this type of data, as it gives a working approach for reducing the amount of data needed for training.
I will briefly go over the 3D mesh/pointcloud semantic segmentation task, and active learning, so that it's clear for the audience not familiar with these concepts. Then I will present the PointNet++ model architecture and Monte-Carlo Dropout approach, that are specifically used in the experiments. And finally, I will share the experiment results with the audience.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Dmitry-Korobchenko.pptx
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Tiling & Parallel Processing of Large Images
During this session, we will review the benefits of processing large imageries by tiles, review use cases, and later combination of results.
As previously mentioned we will review the benefits of tile level processing for large imageries, going further into some use cases seen in data analysis, software engineering, and ML models(such as classification and segmentation). We will review how the tile data was later combined for each use case separately, and what were the benefits we saw from adopting this approach.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Anush-Tosunyan.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Use AutoML to Create High-Quality Models
AWS provides a range of AutoML solutions for all levels of expertise. In this session, we will cover AutoGluon, a library for ML practitioners seeking an open source solution, and Amazon SageMaker tool for data scientists who prefer a fully-managed service. Developers or business users without ML experience can take advantage of ready-made solutions for use cases such as computer vision, demand forecasting, intelligent search, and industrial and healthcare verticals.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Aleksandr-Patrushev.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
The Structure and Interpretation of ML Metadata
This talk is targeted both for ML researchers and engineers working on the ML infrastructure. The metadata generated at almost every step of the ML pipeline connects and enables the reproducible and explainable ML infrastructure. During this talk, we will go through how and where metadata is generated in ML infrastructure. The types of the metadata. What's the next generation ML infra stack to leverage the metadata and help build reproducibility, explainability, and governance into your ML systems?
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Gevorg-Soghomonyan.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Streamlit: A Faster Way to Build and Share Data Apps
Poor tooling slows down data science and machine learning projects.
Streamlit is a fast way to build and share data apps. It is able to turn data scripts into shareable web apps with minimal effort. Let's hear Karen Javadyanl introduce Streamlit: the fastest way to build and share data apps as Python scripts.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Karen-Javadyan.pptx
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Best Practices for Coding in ML/DS - Techniques to Construct your Project
Many engineers, particularly those in Data Science, do not focus on writing better code, which their coworkers will love. This is bad!
Writing cleaner code, and using appropriate tools for experiment logging reduces the time of debugging and the effort spent on the project in the long term. Consequently, the code becomes readable and onboarding new engineers on the project becomes easier.
By the end of the lecture, attendees will have learned about the importance of having a clean code in the ML project. They will have developed intuition about wiring readable and understandable code and will have acquired knowledge about the general design of a good codebase, and some tools that will help engineers log experiments for a cleaner environment.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Mher-Khachatryan.pptx
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Cifar-10 Exploratory Data Analysis
Image classification datasets are completed from the analysis point of view, taking into account the complicated structure of images. However, the understanding of the dataset descriptors at the high level can add debugging facilities and, in early stage, predict the quality of the classification model. During this session, we will visually analyze one of the challenging SOTA datasets like Cifar-10.
The datasets in AI used to contain 1000+ images. Images are matrices, and the handling of available features, missing features that can lead to AI model overfitting or underfitting. Based on visualization will predict whether we can reduce the dataset and come up with a smaller set and predict its impact on the final AI model.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Anna_Shahinyan_16_9_pydata_2022.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Recommendation Systems in Market Research
Gathering opinion data at scale in market research assumes a platform where many users complete surveys from many different providers. As a result, the problem of matching surveys with users arises. Taking into account specifics of the market research industry, recommendation systems, multicriteria optimization, and regression models become a crucial part of efficient user-survey matching at scale.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Davit-Abgaryan.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
Modern Data Stack: Optimising and Scaling Data in a Tech Company
This talk is about a new approach to data integration that DataOps is enabling in tech companies (with Sololearn practical example) to save engineering time, allowing engineers and analysts to pursue higher-value activities, explaining why every tech organisation should have a Data + Analytics Engineering (DataOps) department.
It answers two main questions: Why should we care about the adoption of data + analytics engineering? What are the steps and processes to start the journey to full adoption?
This talk will look at themes around that journey: metadata, tools, & organisation action points to paint a picture of what the next phase of the journey looks like. We will go through the modern data stack. We will talk about its architecture, which tools fit where, and how to organise teams to support it. We’ll also touch on the challenges in operationalising warehouses and discuss future technology advancements that could unlock the warehouse to become the platform for business needs and operations.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Nacho-Aranguren.pptx
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
ML Platform for Insurance Conglomerate
The guide through modern ML platform development for the insurance sector and challenges around.
The talk will be focused on a descriptive guide on how Grid Dynamics was building an ML platform for one of the major US insurance companies, challenges that we have faced, and business benefits clients gained with the new cloud platform.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Dmitry-Mezhensky.pdf
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
PyTorch Geometric for Graph Neural Nets
In contrast to classical Deep Learning models (such as MLP, CNN, RNN, Transformers), which are usually applied to tensors and sequences, Graph Neural Net (GNN) is a special type of Deep Learning model which works with non-euclidian data structures, such as graphs. Examples of graph analysis tasks where a data-driven approach can help may include 3D mesh processing, molecular analysis, social graphs data mining, and potentially any other task where traditional DL methods are inapplicable.
PyTorch is an industry-standard Deep Learning framework which provides a lot of useful DL operations and utilities. PyTorch Geometric is a library built on top of PyTorch, implementing a set of tools to create and train Graph Neural Networks.
In this talk, I will give a very quick and high-level introduction to GNNs and PyTorch Geometrics.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Dmitry-Korobchenko.pptx
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps
The Dangers of Mindless Forecasting
"Prediction is very difficult, especially if it’s about the future!" This phrase is attributed to Niels Bohr, the Nobel laureate in Physics and father of the atomic model. This quote warns about the unreliability of forecasts without proper testing and about constant changes in the initial assumed conditions.
With modern programming languages and convenient packages that provide ready-made modeling solutions, it is often easy to find a model that fits the past data well; perhaps too well! But does the maximization of metrics justify the means? Should the complex structures of predictions be built on the quicksand of noisy data?
This talk is a laid-back discussion that will be useful for the audience from any background, from beginner to advanced. Aghasi Tavadyan is the founder of Tvyal.com, which translates to "data" from Armenian. You can find more info about him following these websites: tavadyan.com, tvyal.com.
Presentation Slides: https://pydatayerevan.aua.am/files/2022/09/Aghasi-Tavadyan.pptx
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
00:00 Welcome!
00:10 Help us add time stamps or captions to this video! See the description for details.
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: github.com/numfocus/YouTubeVideoTimestamps


