Text Analysis for Social Scientists · A SIM DAD LLC product

Deterministic, provenance-first text analysis.

TASS scores text against open, citable dictionaries and shows its work. Same inputs, byte-identical outputs, every time. Every run writes a manifest recording versions, input hashes, licenses, and citations: the record a methods section and a replication both need.

Ten bundled open dictionaries · statistics benchmarked against R · one-time purchase, no subscription

Two editions, one engine

Free where it counts. Paid where it saves you time.

Both editions run the same deterministic engine and produce the same manifests. The Community Edition is free forever and open source. The GUI Edition adds the no-code desktop workflow, inferential statistics, and publication output.

Open source · Apache-2.0

Community Edition

Free forever

The full scoring engine for people comfortable at a command line, and for AI agents driving it over MCP.

  • Dictionary scoring engine: percent, hits, weighted, and mean metrics; group summaries; trajectories
  • Complete CLI: score, analyze, exemplars, KWIC concordance, ingest, import, author, cite
  • MCP server: the whole engine as tools for your AI agent; TASS itself never calls a model
  • Minimal local web GUI (runs on 127.0.0.1 only)
  • Reproducible .tassproj projects with byte-compare rerun
  • All ten bundled dictionaries and free registry installs
  • Zero runtime dependencies; Node 18 or newer is the only requirement
github.com/SIM-DAD/tass
What ships today

Built for numbers you can defend.

Every claim below describes a shipping capability, documented in the public methods reference and user guide.

01 · Reproducibility

Same inputs, byte-identical outputs

The engine has no clocks and no randomness: fixed column orders, stable sorts, one fixed-precision number formatter. A .tassproj archive seals a complete run; tass project rerun re-executes it and byte-compares every artifact. The answer is REPRODUCED, or a precise report of what moved.

02 · Provenance

Every score traces to open, inspectable dictionaries

Every run writes a manifest: tool and engine versions, SHA-256 of each input, and each lexicon's name, license, license class, and citation. Every bundled dictionary is open and verified for any use. Restricted resources such as LIWC or NRC are never bundled; if you import a copy you have licensed, TASS flags every run that touches it.

03 · Statistics

Benchmarked against R, with a public audit report

Welch and Student t, Mann-Whitney, ANOVA variants, Kruskal-Wallis, post-hoc tests, effect sizes, assumption checks, and Benjamini-Hochberg correction, each validated against R 4.6 golden files with agreement typically to 1e-13. The benchmark report is public, and every comparison can emit a base-R script that reproduces the run independently.

04 · Sentiment

TASS VADER-rules, honestly labeled

A clean-room implementation of the published VADER heuristics (negation, boosters, ALL-CAPS, punctuation, but-clauses, emoticons, idioms) over the bundled MIT-licensed VADER lexicon. It can differ from the reference package in corner cases, so TASS reports it as "TASS VADER-rules compound", never as a canonical VADER score.

05 · AI, outside the pipeline

Driveable by your AI agent over MCP

The MCP server exposes the full workflow (18 tools with statistics and charts installed) to Claude or any MCP client. The division of labor is deliberate: your agent drives, the deterministic engine scores, and TASS itself never calls a model. Agent-produced labels join a corpus with explicit external-classifier provenance.

06 · Trace-back

Every number leads back to quotable text

Exemplars return the top and bottom documents for any category with their matched terms; KWIC gives a keyword-in-context concordance. The validation workspace samples matches for human review and reports precision and coder agreement, so dictionary fit becomes evidence instead of hope.

The bundle

Ten open dictionaries, licenses on the table.

Every bundled dictionary is verified for any use, including commercial. Each run records the name, license, and citation of every lexicon it touched in its provenance manifest. More install free by name from the public registry, integrity-checked against a pinned SHA-256.

Dictionaries bundled with TASS, with coverage and license
Dictionary Coverage License
AFINN-165Sentiment valence (−5 to +5 per word)Apache 2.0
VADER LexiconSentiment valence + intensity (rule model)MIT
labMT 1.0Happiness (hedonometric) valenceCC-BY
Empath194 topical and affective categoriesMIT
SocialSent (2000s)Domain-induced historical sentimentPDDL
Warriner VADValence, arousal, dominance normsCC-BY 4.0
Brysbaert Concreteness NormsConcreteness ratings (~40k words)CC-BY 4.0
Kuperman Age-of-AcquisitionAge-of-acquisition norms (~30k words)CC-BY 4.0
stopwords-iso (English)Function words and stopwordsMIT
TASS Politeness Markers v115 politeness-strategy categoriesCC-BY 4.0

Full citations for every dictionary are printed by tass dicts, written into every manifest, and listed in the methods reference. Bring your own too: author a dictionary in any spreadsheet and import it with row-numbered validation, or import a LIWC-format .dic file you have licensed (flagged academic-only in every run's provenance).

The T-Lex program

Open lexicons, built in the open.

T-Lex is TASS's program of original, citable dictionaries for social science text analysis. Every T-Lex dictionary is released under CC-BY-4.0 (ODC-BY for ratings databases): citation required, any use permitted including commercial, never paywalled, in every edition. Each ships versioned in the public registry, installable by name with tass install, together with the authoring memo and a validation report produced with TASS's own validation workspace.

Politeness Markers v1 · shipping now

Fifteen politeness-strategy categories approximating the Stanford politeness taxonomy at the lexical and phrase level. Bundled with every edition.

Discrete Emotions In development

Anger, fear, joy, sadness, anticipation, trust, surprise, disgust. An original, openly licensed alternative where the incumbent resource is restricted to academic use.

Epistemic Language In development

Hedges, boosters, evidentials, and deontic and epistemic modality in one coherent scheme.

Opinion & Stance In development

Evaluative language calibrated for social science text corpora.

Power & Status In development

Authority, dominance, deference, and political-arena language.

Agency & Communion In development

Agentic and communal language across competence, warmth, and affiliation.

Affective Norms Extension In development

A harmonization and coverage-extension layer over the bundled Warriner VAD norms, targeting modern vocabulary gaps. ODC-BY.

Co-author a dictionary

T-Lex entries are community projects with named, cited co-authors, not anonymous donations. Contributors review and validate word lists against published measurement frameworks, re-score items that need domain expertise, and supplement seed lists from their specialization. Substantive contributors are offered co-authorship on the associated dictionary paper and are named in the dictionary metadata.

Interested? Write to tass@simdadllc.com with your field and the constructs you care about, or read CONTRIBUTING in the registry.

Pricing

Buy it once. Own it. Price on the page.

One-time purchase, no subscription. Every GUI Edition license includes all v1.x updates; later major versions are optional paid upgrades. No cart archaeology, no sales call to learn a number.

Community
$0
Free forever · Apache-2.0
  • Engine, CLI, MCP server
  • Minimal local web GUI
  • Reproducible projects
  • All open dictionaries
Get it on GitHub
Student
$59
One-time · verified student status
  • Full GUI Edition
  • All v1.x updates included
Buy
Academic
$149
One-time · single researcher
  • Full GUI Edition
  • All v1.x updates included
  • Clears a P-card without a signature
Buy
Professional
$399
One-time · no institutional email needed
  • Full GUI Edition
  • All v1.x updates included
  • For consultants, journalists, nonprofits, government
Buy
Lab
$499
One-time · five academic seats
  • Full GUI Edition, five seats
  • All v1.x updates included
  • One grant line for a whole lab
Buy

Site and department licenses start at $1,995; write to tass@simdadllc.com for a quote you can put in a budget before talking to anyone. Every purchase carries an unconditional 30-day money-back guarantee; see the refund policy. Dictionaries are never a paid feature in any edition.

Closed beta · now recruiting

Put TASS to work on your corpus.

Ahead of the public launch, we are opening the GUI Edition to researchers who will use it on their own data and help shape the first release. Participation is free for the duration of the beta, and beta installers ship through this program.

The beta is open to people currently affiliated with a university or an established research institution. Apply from your institutional email address; we use it to confirm eligibility. A short note about your institution, your field, and how you would like to use TASS helps us place you.

Apply from your institutional email