Bob Coecke and Mehrnoosh Sadrzadeh’s Distributional Compositional Categorical (DisCoCat) model reads a sentence’s meaning off its grammar. Grammatical structure and vector-space semantics are made the same compact-closed shape, and a functor carries one into the other.

A grammar, whether a pregroup or equivalently a CCG derivation, is itself a compact-closed category. Each type has adjoints with reductions and playing the role of caps. DisCoCat is a strong monoidal functor from this category of grammatical types into a compact-closed category of meaning spaces, either finite-dimensional vector spaces or relations, so grammatical types become spaces, words become vectors or tensors of the type their grammar assigns, and a grammatical reduction becomes a linear map. A sentence’s meaning is then the tensor contraction obtained by applying the caps along the reduction that proves the string grammatical. Drawn as a string diagram, the pregroup cups wire the argument nouns into the verb tensor.

Definition of DisCoCat

Fix a compact-closed grammar category and a compact-closed semantics category (e.g. ). A DisCoCat model is a strong monoidal functor . For a transitive sentence with noun space , sentence space and verb tensor , the meaning is the contraction

where is the cap image of the pregroup reduction.

Remark on language as an enriched category (Bradley)

Tai-Danae Bradley (with Terilla and Vlassopoulos) offers a different categorification of “meaning from usage”. Rather than a functor into vector spaces, this takes an enriched category whose objects are strings/tokens and whose hom-structure records next-token-prediction probabilities. The enriching base is the unit interval as a monoidal poset, so a hom is a probability that one expression extends another. Composition is the chain rule of conditional probabilities, and language itself becomes the category. The compositional content lives in the enrichment instead of in a compact-closed grammar.