
Loading…
Book summary
by Cathy O'Neil
Premium summary · Opens in the app · 30 min read
Data science became one of the most hyped terms of the early twenty-first century. Harvard Business Review called it the sexiest job of the decade. Companies scrambled to hire data scientists, often without knowing exactly what they were hiring for. Universities launched programs. Conferences multiplied. The term appeared everywhere and meant almost anything.
**Straight Talk from the Frontline**
By Cathy O'Neil and Rachel Schutt
*Estimated Reading Time: 45 minutes*
**What You'll Learn**
The real practice of data science, not the hype. You will learn how data scientists think, the algorithms they use, the pitfalls that destroy models, and the human skills that separate effective practitioners from technically proficient failures. This book covers statistical thinking, machine learning algorithms, data wrangling, feature engineering, causality, big data engineering, model evaluation, data leakage, visualization, and the ethical responsibilities that come with working on real-world problems.
**Who This Book Is For**
Anyone who wants to understand what data science actually involves. Whether you are a software engineer moving toward data work, a statistician curious about industry practice, a manager hiring data scientists, or a student considering the field, this book gives you an honest, grounded introduction to the discipline. No advanced mathematics required, but a willingness to think carefully about data and uncertainty is essential.
Data science became one of the most hyped terms of the early twenty-first century. Harvard Business Review called it the sexiest job of the decade. Companies scrambled to hire data scientists, often without knowing exactly what they were hiring for. Universities launched programs. Conferences multiplied. The term appeared everywhere and meant almost anything. This book exists because the hype created confusion. People wanted to know what data science actually is, what data scientists actually do, and how the work differs from statistics, machine learning, or software engineering. The answer, it turns out, is both simple and complex. The simple answer comes from a joke that circulated widely in the field: a data scientist is someone who is better at statistics than any software engineer and better at software engineering than any statistician. The joke captures something real. Data science sits at the intersection of multiple disciplines. It requires technical skill, statistical thinking, and domain knowledge. It demands the ability to write code that works in production and the judgment to know when a model is telling you something meaningful rather than just fitting noise. But the deeper answer is more interesting. Data science is not just a collection of skills. It is a way of approaching problems. It is a process that starts with a question, moves through data collection and cleaning, explores patterns, builds models, evaluates those models honestly, and then deploys them in ways that affect real people. Each step involves judgment. Each step involves choices that can go wrong in subtle ways. The problem this book addresses is that most people learn data science through disconnected tutorials. They learn how to run a linear regression in Python. They learn how to use a machine learning library. They learn how to…
Continue reading in the MinuteRead app
Get the complete 30-minute summary of Doing Data Science
Get the complete summary in the appData science combines statistics, computer science, and domain expertise. It is a distinct discipline, not just a rebran
The data science process is cyclical: question, data collection, cleaning, exploration, modeling, evaluation, deployment
Data is not objective. It reflects choices made during collection and processing. Even large datasets are biased samples
Exploratory data analysis is essential. Always plot your data before modeling.
Data wrangling is often the majority of the work. Master the tools and respect the process.
Feature engineering is where domain expertise becomes concrete. It is often the difference between success and failure.
"Doing Data Science" is a strong fit if you want practical ideas around computer science, technology, programming, especially themes like data science combines statistics, computer science, and domain expertise. it is a distinct discipline, not just a rebran; the data science process is cyclical: question, data collection, cleaning, exploration, modeling, evaluation, deployment. The MinuteRead summary distills these concepts into a focused read, whether you're deciding whether to buy the book or applying its lessons at work.
Motivated to help readers with data scientist (noun): Person who is better at statistics than any software engineer and better at software, Cathy O'Neil wrote “Doing Data Science” to package those ideas for a fast, focused read. In “Doing Data Science”, Cathy O'Neil focuses on data scientist (noun): Person who is better at statistics than any software engineer and better at software. Through “Doing Data Science”, Cathy O'Neil distills the core ideas on computer science into lessons readers can a…
View all summaries by Cathy O'NeilContinue Reading
Access the complete 30-minute summary and thousands more nonfiction books in the MinuteRead app.
Continue reading the complete summary in the MinuteRead app.