SamyakComputer ClassesShakarpur

Course

Data Science

Statistics, Python, SQL and machine learning taught as one connected discipline, with the emphasis on framing a problem correctly and knowing when a result is not real.

  • Duration: 7 months
  • Classroom · Online live
  • Level: intermediate

What you will be able to do

Who this course is for

Syllabus

9 modules · 7 months

  1. Module 1. Statistics that hold up

    • Descriptive statistics and distribution shape
    • Sampling, sampling error and confidence intervals
    • Hypothesis testing and what a p-value does not tell you
    • Correlation, causation and confounding variables
    • Experiment design and reading an A/B test honestly
  2. Module 2. Python for data work

    • Python fundamentals oriented toward data
    • pandas indexing, joining, grouping and reshaping
    • Handling missing data without quietly biasing the result
    • Vectorised thinking with NumPy
    • Reproducible notebooks and project structure
  3. Module 3. SQL and data acquisition

    • Relational modelling and joins at analytical scale
    • Window functions for cohort and time-series work
    • Query performance on large tables
    • APIs, files and combining heterogeneous sources
  4. Module 4. Exploratory analysis

    • Profiling a dataset before modelling it
    • Outlier detection and deciding what to do about them
    • Feature distributions and transformations
    • Visual analysis that reveals rather than decorates
  5. Module 5. Supervised machine learning

    • Regression and classification framing
    • Linear and logistic models, and reading their coefficients
    • Trees, random forests and gradient boosting
    • Train, validation and test discipline, and how leakage happens
    • Cross-validation and hyperparameter search
    • Metric selection and the cost of each error type
  6. Module 6. Unsupervised methods

    • Clustering with k-means and hierarchical methods
    • Dimensionality reduction with PCA
    • Judging whether a cluster is real or an artefact
  7. Module 7. Time series

    • Trend, seasonality and stationarity
    • Baseline forecasts and why they are hard to beat
    • ARIMA-family models
    • Backtesting a forecast honestly
  8. Module 8. Communication and deployment basics

    • Writing up a model for a decision-maker
    • Quantifying and communicating uncertainty
    • Model monitoring and drift
    • Packaging a model behind a simple interface
  9. Module 9. Capstone and interviews

    • Scoping a capstone that can be finished and defended
    • Case-study interview practice
    • Statistics and modelling question drills

Tools and technologies you will use

Projects you will build

Where this course can take you

  • Data Scientist
  • Machine Learning Engineer
  • Quantitative Analyst
  • Senior Data Analyst
  • Research Analyst

Duration, modes and fees

Duration
7 months
Delivery modes
Classroom · Online live
Fees
Share your details for the current fee
Fees vary by batch and delivery mode. Share your details and an advisor will confirm the current fee.

Placement assistance

Every student gets placement assistance — that is what 100% placement assistance means. It is support for all, not a job for all. We do not promise a specific salary, a specific number of interviews, or placement at any named company, and you should be wary of anyone who does.

What is included

  • A place in the monthly placement drive, held every third Saturday
  • The readiness programme every second Saturday — mock interviews and preparation
  • CV review against the specific roles you are targeting
  • Portfolio review, so your project work is presented the way a reviewer will read it
  • Access to the vacancy pool employers send directly to the Samyak network
  • Guidance on which roles realistically fit your background and which do not
  • A place in the next drive, with coaching, if you are not selected in this one

What is not included

  • Any guarantee of a job, an interview, or a particular salary
  • Placement at a named or partner company
  • Applying to jobs on your behalf
  • Support before you have completed the course and its project work
  • Visa, relocation or overseas placement assistance

The skill this course is actually about

It is not model training. Fitting a gradient boosting model is four lines of code and any tutorial will show you.

The skill is knowing whether the number that comes out means anything — whether the training set leaked information from the future, whether the test split respected time order, whether the metric you chose rewards the behaviour you want, whether the effect survives a different sample. Most models that fail in production were technically well-fitted and conceptually wrong.

So statistics comes first here, before any machine learning. Understanding sampling variation and confounding is what lets you look at a 94% accuracy score and ask the right follow-up question.

Baselines, everywhere

Every modelling project in this course requires a baseline first — predict the mean, predict last month, predict the majority class. You are not allowed to report a model result without it.

This is unglamorous and it is the fastest way to develop judgement. A sophisticated model that barely beats “predict last month” is not a success, and noticing that early is what separates useful data science from expensive theatre.

Honest expectations

Seven months gets you genuinely competent at framing problems, building pipelines, modelling carefully and communicating results. It does not make you a research scientist, and it does not substitute for domain knowledge in a specific industry. Many of our learners enter as analysts and move into data science roles from inside a company, which is a well-trodden and realistic path.

Questions

Data Science — frequently asked questions

What is the real difference between data analytics and data science?

Analytics explains what happened and why, using SQL, dashboards and statistics. Data science extends that into prediction and inference — modelling what is likely to happen and quantifying confidence. Analytics roles are more numerous at entry level, so if you are starting from zero we usually suggest analytics first and this course second.

Do I need a mathematics or statistics degree?

No. You need school-level mathematics and willingness to work through the statistics module properly. We teach statistics as applied reasoning rather than proof. What genuinely matters more is scepticism — the instinct to ask whether a result could be an artefact — and that is taught by practice, not by prior credentials.

Will this course get me a data scientist job directly?

It gives you the skills and a defensible portfolio. Whether it gets you the title depends on the market and your background; many people enter as an analyst and move across within a year or two. Anyone promising a data scientist role as an outcome of a course is overselling, and we would rather set the expectation accurately.

How much overlap is there with your AI course?

Some, in the machine learning foundations. This course goes deeper into statistics, inference, experimental design and time series. The AI course goes deeper into deep learning, transformers and generative systems. Choose by the work you want — reasoning from data, or building AI applications.

Is seven months realistic alongside a job?

Yes, and most of our learners do it that way, with weekday evening or weekend batches. Budget around ten hours a week outside class. The statistics and machine learning modules in particular reward practice between sessions rather than passive attendance.

Enquire about Data Science

Three details is all we need. A course advisor will call you back.

By submitting, you agree to be contacted about courses and accept our privacy policy.

Next step

Talk to a course advisor

Tell us what you want to learn and we will help you pick the right course, batch and mode.

Request a callback

Three details is all we need. A course advisor will call you back.

By submitting, you agree to be contacted about courses and accept our privacy policy.