Case Study · Data Analytics & Business Intelligence


Everyone calls it a tech corridor. The data says life sciences.


An end-to-end analysis of 1,348 Greater Boston companies — where a fragmented industry taxonomy was hiding the market's real leader.

Python Pandas Exploratory analysis Feature engineering Tableau
212 · LIFE SCIENCES 204 · IT SERVICES

1,348
Companies across 11 cities
103 8
Industry labels consolidated into real sectors
1
Leading sector corrected
100%
Reproducible — notebook public, dashboard live


The problem


Three simple questions that the data answered incorrectly


Anyone weighing entry, expansion, or investment in Greater Boston's western corridor needs three things first: where the companies are, what they do, and how big they are.

Those questions sound trivial. They aren't. The dataset that carries the answers — 1,348 records, 19 fields, exported from a commercial provider — encodes its own assumptions about how the world is organized. One of them was wrong in a way that would have led a decision-maker to misread the entire market.

How big

Characterize size against a severely skewed distribution, where the average describes almost nobody.

So what

Deliver it so a non-technical decision-maker gets the finding in three seconds.



The finding


Six labels, one sector, and a leader nobody could see


The industry field held 103 categories. At face value, Information Technology & Services led the market with 204 companies.

But six of those categories weren't separate industries. Hospital care, medical devices, pharmaceuticals, biotechnology, mental health care, wellness — facets of a single sector, split apart by the source system's classification scheme. Individually, each looked modest. Together, they outrank IT.

Leading sector: Information Technology & Services

As the source labeled it. Healthcare sits in six pieces, none large enough to lead. A report built on these categories would have named the wrong industry — confidently, and in good faith.


Phase 01 · Python



Finding the truth before anyone sees a chart


Analysis is a sequence of judgment calls made before a single visual exists. These were the ones that mattered.

01

Find what the data can't answer

The brief included a fourth question: which job positions are advertised? The dataset had no such field — it was a company export, not a jobs feed. I caught it in the first ten minutes and renegotiated the scope. Knowing the boundary of your evidence isn't a limitation on the analysis. It is the analysis.

02

Distrust a field that reports itself as clean

Completeness checks said the location data was 100% populated. True — and misleading. Null-counts detect missing values, not meaningless ones. The State field held a single value across all 1,348 records: a filter applied upstream, silently scoping every conclusion I was about to draw.

No variance means no information — but it carried provenance. I extracted the fact, stated it as a limitation, and dropped the field.

03

Engineer the feature that changes the answer

103 labels is a high-cardinality problem with a hidden structure. I mapped the fragmented healthcare and education categories into coherent sectors — reuniting what the taxonomy had split.

Healthcare & Life Sciences: 212 companies. IT Services: 204. The market leader changed.

04

Refuse to delete the outliers

Employee counts ran from 1 to 364,000 against a median of 62. A reflexive cleaning step would have stripped the extremes — which turned out to be TJX, Thermo Fisher, Harvard, MIT, and Beth Israel Lahey. The most significant records in the set.

An outlier is a statistical position, not a verdict of error. I kept them and changed the statistic instead: median and interquartile range, never the mean.

05

Catch the artifact before it becomes a headline

A revenue-per-employee metric ranked a one-person dog-walking business as the region's most capital-efficient company, at $49M per head. That's not an insight — it's a near-zero denominator.

Filtering to credible headcounts surfaced the real pattern: biotech, energy, IT, and finance lead on revenue efficiency. Confident nonsense is worse than no answer, because it gets acted on.

06

Read charts against each other

By total revenue, Cambridge ranked first and Framingham third — implying they're comparable. They aren't. Cambridge's weight spreads across research, energy, biotech, healthcare and education. Framingham's rests on two retailers; remove one company and its ranking collapses.

A dashboard reporting only the ranking would have invited a strategic error.

Every decision is documented and reproducible in the public notebook.



Phase 02 · Tableau



The dashboard:


I designed backwards from one sentence: Healthcare & Life Sciences — not IT — lead a Cambridge-anchored corridor. Every element had to earn its place by advancing that claim.


Live & interactive on Tableau Public Open in Tableau Public →



Outcomes


What the market actually looks like

212

Healthcare & Life Sciences leads — 16% of the market, visible only after correcting the taxonomy that made IT appear to lead.

35%

One city holds a third of the market. Cambridge alone; 72% sits in the top four. This is a cluster, not a region.

62

Median employees per company — against a mean of 1,062. A small-and-mid-cap market the average completely misrepresents.

Concentration isn't uniform in kind. Cambridge is diversified; a similar-looking city rests on a single employer.

The data's categories:

A classification scheme encodes someone else's assumptions about what belongs together. Accepting them unexamined would have produced a technically accurate, entirely wrong conclusion.

The outliers were Harvard, MIT, and a Fortune 100 retailer. The right response to a skewed distribution isn't to remove the tail.

The top of any per-unit ranking is dominated by small-denominator noise. Ranking without interrogating the denominator produces confident nonsense

Two of my most sophisticated visuals never made the final canvas. They were interesting but not part of the message. The hardest editing is cutting your own best work.

Python Pandas / NumPy Exploratory data analysis Data quality auditing Feature engineering Statistical reasoning Tableau Dashboard design Data storytelling Executive communication

"Data doesn't speak for itself. It speaks in the language of whoever structured it. The job is knowing when that structure is hiding flaws, therefore reveling the real story."

Phillippe Jardim · Project Leader & Systems Engineer

Previous
Previous

Artificial Intelligence