CSV / Data Upload Analyzer

Upload a CSV or TSV file to inspect columns, calculate descriptive statistics, view distributions, flag potential outliers and check missing values. Numeric columns also get a Q-Q plot and a normality check when the sample is suitable.

Browser-side analysis Histogram + box plot IQR outlier flags Categorical frequencies No sign-up

Upload Your Data

CSV and TSV are supported. Excel files are not parsed on this page.

Your dataset stays in the analyzer running in this browser.This page's analyzer code reads and calculates from the file locally. It does not send file contents to a server. Use Clear Data to remove the current dataset from the page.

Drop a CSV or TSV here, or click to choose a file

Quoted fields, embedded commas, UTF-8 text and common line endings are handled.

Limit: 8 MB and 100,000 data rows per analysis.
Parsing settings
Empty cells are always treated as missing. Edit this list if “NA” or “null” are legitimate text values.

Dataset Overview

Uploaded data

Columns and Detected Types

Automatic type detection can be wrong. Change a type before interpreting the results.

ColumnDetectedUse asValidMissing

Data Preview

First 12 rows only. The full parsed column is used for calculations.

Analyze a Column

Select one variable to see its summary and diagnostics.

What Does the CSV Data Analyzer Calculate?

A CSV data analyzer reads a tabular file and summarizes each variable without requiring a statistics package. This page detects common column types, reports missing values, calculates descriptive statistics for numeric data and builds frequency tables for categorical data. Numeric columns can also be inspected with a histogram, box plot, potential-outlier list and normal Q-Q plot.

The tool is designed for exploratory data analysis. It helps you see what is in a file before you choose a formal statistical method. It does not determine whether a value is scientifically correct, whether an association is causal or whether a flagged observation should be deleted.

Descriptive Statistics Explained

For a numeric column, the analyzer removes configured missing values from the calculation and reports how many were excluded. It then calculates the mean, median, minimum, maximum, range, quartiles, interquartile range, sample variance, sample standard deviation, standard error, skewness and excess kurtosis when the sample size allows those quantities to be estimated.

Mean: x̄ = Σx / n
Sample variance: s² = Σ(x − x̄)² / (n − 1)
Sample standard deviation: s = √s²
Range: max − min
IQR: Q3 − Q1

The quartiles on this page use linear interpolation with the index (n − 1)p, sometimes called the R-7 percentile convention. Different software can use different quartile definitions, so Q1 and Q3 can differ slightly between programs even when the same values are supplied.

For a manual-input alternative, use the descriptive statistics calculator. The descriptive statistics guide explains the ideas behind these summaries.

Understanding Detected Column Types

Column type detection is a starting point, not a statement about what a variable means. The analyzer samples nonmissing values and checks whether most values can be parsed consistently as numbers, dates or simple boolean labels. Anything else is treated as categorical text, while mixed patterns are flagged for review.

A column of student numbers, invoice numbers or customer IDs can contain only digits and still be categorical in practice. Averaging an identifier has no useful interpretation. For that reason, the tool marks some unique integer columns as possible IDs and lets you override the detected type. The same principle applies in the other direction: a numeric measurement can be stored as text because a few rows include units such as 12 kg or a note such as pending.

Changing the type does not rewrite the source data. It only changes how the selected column is interpreted for this analysis. If you choose Numeric, values that cannot be parsed as finite numbers are counted as invalid nonmissing entries and excluded from numeric calculations rather than being silently converted.

How Outlier Detection Works

The outlier panel uses the 1.5 × IQR rule. It calculates a lower fence of Q1 − 1.5 × IQR and an upper fence of Q3 + 1.5 × IQR. Observations outside those fences are labeled potential outliers.

A potential outlier is not automatically an error. It can be a valid extreme observation, a recording mistake or an important part of the phenomenon being studied. Review the original record and subject-matter context before changing or removing it.

The box plot uses the same quartile convention as the numeric summary and the outlier check. Its whiskers stop at the most extreme observed values that remain inside the IQR fences. Read more about the interquartile range and outliers in descriptive statistics.

Reading the Histogram and Box Plot Together

The histogram groups numeric observations into intervals, which makes the overall shape easier to inspect. Look for a central cluster, long tails, multiple peaks, gaps and isolated bars. The automatic setting uses the Freedman-Diaconis rule when the IQR is positive because it adapts bin width to both sample size and spread. If that rule cannot be used, the analyzer falls back to Sturges' rule.

No single histogram is definitive. A narrow bin width can expose small bumps that disappear with wider bins, while wide bins can hide local structure. The bin selector is included so you can check whether a visual pattern is stable rather than trusting one automatic setting.

The box plot answers a different question. It emphasizes the median, middle 50% of the data, whiskers and observations beyond the IQR fences. A box plot can show asymmetry and unusually distant observations quickly, but it does not reveal multiple peaks as clearly as a histogram. Use both when distribution shape matters. The histogram maker and box plot generator are available for focused chart work.

How the Normality Check Works

Normality is not judged from a p-value alone. The analyzer shows a normal Q-Q plot so you can inspect how closely the ordered observations follow a straight reference line. For nonconstant numeric columns with at least 20 valid observations, it also reports the Jarque-Bera test, which checks whether sample skewness and kurtosis differ from values expected under a normal model.

A small p-value is evidence against exact normality under the test. A larger p-value does not prove that the population is normal. Small samples can miss meaningful departures, while very large samples can make small deviations statistically detectable. Use the plot, the test and the purpose of the planned analysis together.

If you need a dedicated Shapiro-Wilk procedure, see the Shapiro-Wilk test guide and the broader normality tests overview.

How Missing Values Are Handled

Empty cells are missing by default. The parsing settings also contain an editable list of text markers such as NA, N/A and null. Because those strings can be legitimate category labels in some datasets, the list is visible and can be changed before analysis.

Univariate numeric summaries omit missing cells and report the number omitted. A value that is present but cannot be parsed as a number is not silently converted to zero. The column table reports valid and missing counts, and the data-quality panel warns when a selected numeric type still contains nonnumeric values.

CSV Formatting Guide

The parser supports comma, tab and semicolon delimiters. Auto detection compares the first part of the file and chooses the delimiter that produces the most consistent field count. Quoted fields can contain delimiters, escaped quotation marks and line breaks. That matters for ordinary CSV exports such as a comments field containing "London, UK".

Input featureSupportedNotes
CSVYesComma delimiter or auto detected.
TSVYesTab-delimited text is supported.
Semicolon-delimited textYesUseful for some regional spreadsheet exports.
Quoted delimitersYesFields such as "Smith, Jane" stay together.
UTF-8 textYesModern browser text decoding is used.
XLSX / XLSNoExport the worksheet as CSV or TSV first.

Privacy and Data Processing

File reading, parsing, calculations and chart generation in this analyzer are implemented in browser JavaScript. The analyzer code does not submit the dataset to StatisticsFundamentals.com for processing. It also does not place the parsed dataset in local storage. Clicking Clear Data removes the current in-memory references and resets the file input and results.

This describes the analyzer code on the page, not the security policy of the device, browser extensions or network environment. Avoid opening confidential datasets on a computer you do not control.

Why Results Can Differ From Excel, R or Other Tools

Two programs can analyze the same values and still return slightly different quartiles, variance estimates or outlier fences. The difference is often methodological rather than a calculation error. Percentile software uses several accepted interpolation conventions, and variance can be calculated with either n or n − 1 in the denominator depending on whether a population or sample quantity is intended.

This analyzer uses sample variance and sample standard deviation with n − 1. Quartiles use the R-7 linear interpolation convention. Empty cells and the visible missing-value markers are excluded from univariate summaries. Displayed numbers are rounded for readability, while calculations retain JavaScript's full numeric precision.

Parsing can also create differences. A cell containing a currency symbol, a percent sign or a thousands separator is not automatically stripped and guessed as a number. That conservative behavior reduces silent data changes, but it means you may need to clean a column before analysis. If results disagree with another program, compare the raw values, missing-value rules, variance definition and quartile method before assuming one result is wrong.

Scope and Limits of This Analyzer

This page performs univariate exploratory analysis. It is intended to answer questions such as how a variable is distributed, how much data are missing, what its central and spread statistics are and whether any observations fall beyond a common descriptive outlier rule. It does not fit regression models, compare experimental groups, calculate causal effects or choose a hypothesis test for you.

The 8 MB and 100,000-row limits are deliberate browser safeguards. Summary statistics and histogram counts use the full selected numeric column inside those limits. A Q-Q plot with more than 1,000 valid values displays 1,000 evenly spaced ordered points so the page does not create tens of thousands of SVG elements. The normality test still uses the full valid column. The box plot can draw up to 300 individual outlier markers while the reported outlier count uses all valid values.

Exploratory statistics should be checked against the way the data were collected. A clean-looking distribution cannot establish measurement quality, independence, representativeness or causality. Those questions depend on study design and subject knowledge, not on a CSV summary alone.

Common Data Analysis Mistakes

MistakeWhy it matters
Treating missing cells as zeroZero is an observed value. Missing means no usable value was recorded.
Deleting every flagged outlierThe IQR rule flags unusual values; it does not determine whether they are wrong.
Using only the meanSkewed data can have a mean that does not describe a typical observation well.
Calling an ID a measurementA numeric identifier can produce a mean and SD that have no substantive meaning.
Reading p > .05 as proof of normalityA test can fail to reject normality without establishing that the model is exactly true.
Ignoring the histogram bin widthChanging bins can make the same dataset look smoother or more irregular.
Assuming correlation means causationExploratory patterns do not identify a causal mechanism by themselves.

Related Statistics Tools and Guides

Frequently Asked Questions

Upload the CSV, review the detected column types, then choose a column in the analysis panel. Numeric columns receive descriptive statistics and distribution diagnostics. Categorical columns receive frequency counts and percentages.

For numeric data, this analyzer calculates count, missing values, mean, median, mode when repeated values exist, minimum, maximum, range, Q1, Q3, IQR, sample variance, sample standard deviation, sum, standard error, skewness and excess kurtosis where sample size permits.

The analyzer code processes file contents in the browser and does not send the dataset to a server for calculation. It does not save the parsed dataset in local storage. Clear Data resets the current in-memory analysis.

Not on this version of the analyzer. Export the worksheet as CSV or TSV, then upload that file. The page does not claim XLSX or XLS support.

The default rule flags values below Q1 − 1.5 × IQR or above Q3 + 1.5 × IQR. These are descriptive flags, not automatic errors, and the analyzer does not delete them.

For suitable numeric samples, the page reports the Jarque-Bera test and pairs it with a normal Q-Q plot. The test uses skewness and kurtosis. It is not run for very small or constant samples.

No. It means the test did not find enough evidence to reject the specified normal model at that significance threshold. It does not prove exact normality. Sample size and the Q-Q plot also matter.

Type detection is conservative when values are mixed. Symbols, units, thousands separators, unexpected text or inconsistent date formats can cause a numeric-looking column to be classified as categorical or mixed. Use the type dropdown only after checking the raw values.

Common reasons include different missing-value rules, sample versus population variance, quartile conventions, rounding and parser behavior. This page labels its variance denominator and quartile method so you can compare methods.

This implementation limits each analysis to 8 MB and 100,000 data rows. The cap reduces the chance of a browser tab becoming unresponsive, especially on phones and older computers.