What Does the CSV Data Analyzer Calculate?
A CSV data analyzer reads a tabular file and summarizes each variable without requiring a statistics package. This page detects common column types, reports missing values, calculates descriptive statistics for numeric data and builds frequency tables for categorical data. Numeric columns can also be inspected with a histogram, box plot, potential-outlier list and normal Q-Q plot.
The tool is designed for exploratory data analysis. It helps you see what is in a file before you choose a formal statistical method. It does not determine whether a value is scientifically correct, whether an association is causal or whether a flagged observation should be deleted.
Descriptive Statistics Explained
For a numeric column, the analyzer removes configured missing values from the calculation and reports how many were excluded. It then calculates the mean, median, minimum, maximum, range, quartiles, interquartile range, sample variance, sample standard deviation, standard error, skewness and excess kurtosis when the sample size allows those quantities to be estimated.
The quartiles on this page use linear interpolation with the index (n − 1)p, sometimes called the R-7 percentile convention. Different software can use different quartile definitions, so Q1 and Q3 can differ slightly between programs even when the same values are supplied.
For a manual-input alternative, use the descriptive statistics calculator. The descriptive statistics guide explains the ideas behind these summaries.
Understanding Detected Column Types
Column type detection is a starting point, not a statement about what a variable means. The analyzer samples nonmissing values and checks whether most values can be parsed consistently as numbers, dates or simple boolean labels. Anything else is treated as categorical text, while mixed patterns are flagged for review.
A column of student numbers, invoice numbers or customer IDs can contain only digits and still be categorical in practice. Averaging an identifier has no useful interpretation. For that reason, the tool marks some unique integer columns as possible IDs and lets you override the detected type. The same principle applies in the other direction: a numeric measurement can be stored as text because a few rows include units such as 12 kg or a note such as pending.
Changing the type does not rewrite the source data. It only changes how the selected column is interpreted for this analysis. If you choose Numeric, values that cannot be parsed as finite numbers are counted as invalid nonmissing entries and excluded from numeric calculations rather than being silently converted.
How Outlier Detection Works
The outlier panel uses the 1.5 × IQR rule. It calculates a lower fence of Q1 − 1.5 × IQR and an upper fence of Q3 + 1.5 × IQR. Observations outside those fences are labeled potential outliers.
A potential outlier is not automatically an error. It can be a valid extreme observation, a recording mistake or an important part of the phenomenon being studied. Review the original record and subject-matter context before changing or removing it.
The box plot uses the same quartile convention as the numeric summary and the outlier check. Its whiskers stop at the most extreme observed values that remain inside the IQR fences. Read more about the interquartile range and outliers in descriptive statistics.
Reading the Histogram and Box Plot Together
The histogram groups numeric observations into intervals, which makes the overall shape easier to inspect. Look for a central cluster, long tails, multiple peaks, gaps and isolated bars. The automatic setting uses the Freedman-Diaconis rule when the IQR is positive because it adapts bin width to both sample size and spread. If that rule cannot be used, the analyzer falls back to Sturges' rule.
No single histogram is definitive. A narrow bin width can expose small bumps that disappear with wider bins, while wide bins can hide local structure. The bin selector is included so you can check whether a visual pattern is stable rather than trusting one automatic setting.
The box plot answers a different question. It emphasizes the median, middle 50% of the data, whiskers and observations beyond the IQR fences. A box plot can show asymmetry and unusually distant observations quickly, but it does not reveal multiple peaks as clearly as a histogram. Use both when distribution shape matters. The histogram maker and box plot generator are available for focused chart work.
How the Normality Check Works
Normality is not judged from a p-value alone. The analyzer shows a normal Q-Q plot so you can inspect how closely the ordered observations follow a straight reference line. For nonconstant numeric columns with at least 20 valid observations, it also reports the Jarque-Bera test, which checks whether sample skewness and kurtosis differ from values expected under a normal model.
A small p-value is evidence against exact normality under the test. A larger p-value does not prove that the population is normal. Small samples can miss meaningful departures, while very large samples can make small deviations statistically detectable. Use the plot, the test and the purpose of the planned analysis together.
If you need a dedicated Shapiro-Wilk procedure, see the Shapiro-Wilk test guide and the broader normality tests overview.
How Missing Values Are Handled
Empty cells are missing by default. The parsing settings also contain an editable list of text markers such as NA, N/A and null. Because those strings can be legitimate category labels in some datasets, the list is visible and can be changed before analysis.
Univariate numeric summaries omit missing cells and report the number omitted. A value that is present but cannot be parsed as a number is not silently converted to zero. The column table reports valid and missing counts, and the data-quality panel warns when a selected numeric type still contains nonnumeric values.
CSV Formatting Guide
The parser supports comma, tab and semicolon delimiters. Auto detection compares the first part of the file and chooses the delimiter that produces the most consistent field count. Quoted fields can contain delimiters, escaped quotation marks and line breaks. That matters for ordinary CSV exports such as a comments field containing "London, UK".
| Input feature | Supported | Notes |
|---|---|---|
| CSV | Yes | Comma delimiter or auto detected. |
| TSV | Yes | Tab-delimited text is supported. |
| Semicolon-delimited text | Yes | Useful for some regional spreadsheet exports. |
| Quoted delimiters | Yes | Fields such as "Smith, Jane" stay together. |
| UTF-8 text | Yes | Modern browser text decoding is used. |
| XLSX / XLS | No | Export the worksheet as CSV or TSV first. |
Privacy and Data Processing
File reading, parsing, calculations and chart generation in this analyzer are implemented in browser JavaScript. The analyzer code does not submit the dataset to StatisticsFundamentals.com for processing. It also does not place the parsed dataset in local storage. Clicking Clear Data removes the current in-memory references and resets the file input and results.
This describes the analyzer code on the page, not the security policy of the device, browser extensions or network environment. Avoid opening confidential datasets on a computer you do not control.
Why Results Can Differ From Excel, R or Other Tools
Two programs can analyze the same values and still return slightly different quartiles, variance estimates or outlier fences. The difference is often methodological rather than a calculation error. Percentile software uses several accepted interpolation conventions, and variance can be calculated with either n or n − 1 in the denominator depending on whether a population or sample quantity is intended.
This analyzer uses sample variance and sample standard deviation with n − 1. Quartiles use the R-7 linear interpolation convention. Empty cells and the visible missing-value markers are excluded from univariate summaries. Displayed numbers are rounded for readability, while calculations retain JavaScript's full numeric precision.
Parsing can also create differences. A cell containing a currency symbol, a percent sign or a thousands separator is not automatically stripped and guessed as a number. That conservative behavior reduces silent data changes, but it means you may need to clean a column before analysis. If results disagree with another program, compare the raw values, missing-value rules, variance definition and quartile method before assuming one result is wrong.
Scope and Limits of This Analyzer
This page performs univariate exploratory analysis. It is intended to answer questions such as how a variable is distributed, how much data are missing, what its central and spread statistics are and whether any observations fall beyond a common descriptive outlier rule. It does not fit regression models, compare experimental groups, calculate causal effects or choose a hypothesis test for you.
The 8 MB and 100,000-row limits are deliberate browser safeguards. Summary statistics and histogram counts use the full selected numeric column inside those limits. A Q-Q plot with more than 1,000 valid values displays 1,000 evenly spaced ordered points so the page does not create tens of thousands of SVG elements. The normality test still uses the full valid column. The box plot can draw up to 300 individual outlier markers while the reported outlier count uses all valid values.
Exploratory statistics should be checked against the way the data were collected. A clean-looking distribution cannot establish measurement quality, independence, representativeness or causality. Those questions depend on study design and subject knowledge, not on a CSV summary alone.
Common Data Analysis Mistakes
| Mistake | Why it matters |
|---|---|
| Treating missing cells as zero | Zero is an observed value. Missing means no usable value was recorded. |
| Deleting every flagged outlier | The IQR rule flags unusual values; it does not determine whether they are wrong. |
| Using only the mean | Skewed data can have a mean that does not describe a typical observation well. |
| Calling an ID a measurement | A numeric identifier can produce a mean and SD that have no substantive meaning. |
| Reading p > .05 as proof of normality | A test can fail to reject normality without establishing that the model is exactly true. |
| Ignoring the histogram bin width | Changing bins can make the same dataset look smoother or more irregular. |
| Assuming correlation means causation | Exploratory patterns do not identify a causal mechanism by themselves. |
Related Statistics Tools and Guides
Frequently Asked Questions
Upload the CSV, review the detected column types, then choose a column in the analysis panel. Numeric columns receive descriptive statistics and distribution diagnostics. Categorical columns receive frequency counts and percentages.
For numeric data, this analyzer calculates count, missing values, mean, median, mode when repeated values exist, minimum, maximum, range, Q1, Q3, IQR, sample variance, sample standard deviation, sum, standard error, skewness and excess kurtosis where sample size permits.
The analyzer code processes file contents in the browser and does not send the dataset to a server for calculation. It does not save the parsed dataset in local storage. Clear Data resets the current in-memory analysis.
Not on this version of the analyzer. Export the worksheet as CSV or TSV, then upload that file. The page does not claim XLSX or XLS support.
The default rule flags values below Q1 − 1.5 × IQR or above Q3 + 1.5 × IQR. These are descriptive flags, not automatic errors, and the analyzer does not delete them.
For suitable numeric samples, the page reports the Jarque-Bera test and pairs it with a normal Q-Q plot. The test uses skewness and kurtosis. It is not run for very small or constant samples.
No. It means the test did not find enough evidence to reject the specified normal model at that significance threshold. It does not prove exact normality. Sample size and the Q-Q plot also matter.
Type detection is conservative when values are mixed. Symbols, units, thousands separators, unexpected text or inconsistent date formats can cause a numeric-looking column to be classified as categorical or mixed. Use the type dropdown only after checking the raw values.
Common reasons include different missing-value rules, sample versus population variance, quartile conventions, rounding and parser behavior. This page labels its variance denominator and quartile method so you can compare methods.
This implementation limits each analysis to 8 MB and 100,000 data rows. The cap reduces the chance of a browser tab becoming unresponsive, especially on phones and older computers.