Skip to content
CAAIL

Linear & Regularized Models

This page describes the Linear & Regularized Models row of the Papers.md matrix: ordinary and penalized linear and logistic regression (LASSO, ridge, elastic net) and linear additive scoring models. The row’s authoritative scope is its Taxonomy.md definition; this page synthesizes what currently sits in it.

Scope boundary

The taxonomy separates this row from its three nearest neighbours by mechanism rather than by performance: no kernel or margin, which excludes SVM; no trees, bagging or boosting, which excludes Ensemble Learning; and no restriction to spectral latent-variable projection, which distinguishes it from Chemometrics, where PLS is a linear model but a particular one applied to a particular data type.

What the row is for is the more useful framing. These models are chosen when the feature set is modest and someone needs to know which features drive the prediction, because the fitted coefficients name them and an L1 penalty selects them outright. That is what most of the references below are here for: they report feature importance or propensity as a result in its own right rather than only an accuracy number. The exceptions are the comparative studies, where a penalized regressor is one entrant among many and the finding is about method choice. That is a recurring need in cell-ag work, where a model that says “carbohydrate and targeted moisture content dominate texture” changes a formulation decision in a way that a more accurate but opaque model does not.

Cellular Engineering

  • #266 (Wang et al. 2023, Journal of Agricultural and Food Chemistry, Nanjing Agricultural University): predicts the proliferation and differentiation potency of porcine muscle stem cells from cell morphology, aimed squarely at the batch-to-batch quality variation that makes cultured-meat cell production inconsistent. pMuSCs were sorted (CD31−, CD45−, CD56+, CD29+) from three pigs, cultured across passages 5 to 9 in three lots, with proliferation scored as growth rate and differentiation as the average myosin-heavy-chain stained area after five days. The result that matters methodologically is about when to look rather than which model: predictions built on the 36-hour and 60-hour morphological profiles were the best, reaching R² = 0.95 for proliferation and R² = 0.74 for differentiation, and the paper argues that accumulating time-course information about morphological heterogeneity in the population is what makes potency predictable at all. One of the few references in the whole matrix working directly on a livestock cell line intended for cultivated meat.
  • #122 PreciCE (Magnusson et al. 2024, bioRxiv): identifies the smallest set of transcription factors to perturb to move a cell from a starting state to a target state, posed as the recovery of a sparse perturbation vector. The row is earned by the linear case, which the methods set out as such: where the learned map is linear, the perturbation is recovered by “regularized least square regression (lasso)”, with L1 regularization carrying the constraint that only k genes may be non-zero, solved by coordinate descent with the penalty swept across a wide range so solutions at the wanted sparsity come back. The sparsity is the biologically load-bearing part rather than a modelling convenience, since a protocol that perturbs three genes can be run and one that perturbs three hundred cannot. Also in Deep Learning, which covers the neural form of the same map.

Bioprocess & Scale-Up

  • #32 (Roell et al. 2022, Biochemical Engineering Journal): one of the seven algorithm families benchmarked on Clostridium carboxidivorans syngas fermentation, in a comparison whose design decision matters more than its winner: training runs are separated from test runs by fermentation condition, so no condition appears in both. Described in K-Nearest Neighbors; also in SVM, Ensemble Learning and Comparative Studies.
  • #208 (Xu et al. 2025, Biotechnology Progress): the clearest statement in the matrix of why this row exists. Lasso is chosen for its built-in regularization, and Ridge and Elastic Net because they “effectively handle multicollinearity”, which is exactly what full-spectrum Raman and capacitance data produce, with partial least squares carried as the field’s default baseline. Described in Chemometrics; also in SVM, Ensemble Learning and Comparative Studies.
  • #253 (Park et al. 2023, Biotechnology and Bioengineering): the linear and penalized regressors among the twelve classical approaches compared against four deep ones for forecasting fed-batch CHO-K1 profiles. Described in Comparative Studies.

Scaffolding

  • #171 (Kircali Ata et al. 2023, Foods): predicts the hardness and chewiness of plant-based meat analogs from the proximate composition of the raw materials (protein, fat, carbohydrate, fibre, ash, moisture) plus a targeted moisture content, over data curated from three prior extrusion and mechanical-elongation studies. Ridge is the linear member of a comparison that also includes XGBoost and an MLP, all with built-in feature selection, evaluated leave-one-group-out so a whole source study is held out at a time rather than random rows. Reported MAPE 22.9% for hardness and 14.5% for chewiness. The paper also examines multicollinearity among the composition features and the linearity of the design, which is the sort of diagnostic that belongs to this row specifically. Code and the curated table at sezinata/FoodML; the dataset is catalogued in Datasets/CrossSpecies.md. Also in Ensemble Learning, and in this row again under Sensory Prediction.

Sensory Prediction

  • #171 (Kircali Ata et al. 2023, Foods): the same texture-prediction work, placed here as well because hardness and chewiness are sensory attributes as much as structural ones. See the description under Scaffolding above.
  • #269 iUmami-SCM (Charoenkwan et al. 2020, Journal of Chemical Information and Modeling): predicts whether a peptide tastes umami from primary sequence alone, using a scoring card method, a linear additive model over propensity scores for the 20 amino acids and the 400 dipeptides. Built on UMP442, 140 umami and 304 non-umami peptides assembled from the literature and BIOPEP-UWM, with bitter peptides used as the negative class and an 80/20 split into UMP-TR and UMP-IND. Reported accuracy 0.865 and MCC 0.679 on the independent set, outperforming the standard ML classifiers it was compared against. The interpretability is the point rather than a side benefit: the propensity scores are analysed directly to characterize which residues and dipeptides carry umami intensity, which is what a formulator can act on. Benchmark data and the scoring-card implementation are catalogued in Datasets/Benchmarks.md. Also in Genetic Algorithms, which is where the evolutionary search producing these propensity scores is described.
  • #339 (Gutiérrez et al. 2018, Nature Communications, IBM Research): predicts how humans will describe the smell of a mono-molecular odorant, over the Dravnieks dataset of 128 molecules rated by 507 experts across 146 verbal descriptors, plus the Keller and Vosshall data of 476 molecules rated by 49 individuals. The move that lifts the result is representational: descriptors are embedded as 300-dimensional fastText word vectors trained on 16 billion words, so a descriptor with little training data borrows structure from semantically nearby ones. That raised the number of descriptors predictable at accuracy above 0.5 to around 70, roughly a tenfold increase over prior work, and the authors argue the semantic distances between descriptors amount to an odour wheel. Relevant to cell-ag sensory work because panel vocabulary, not instrument data, is usually the scarce resource.

Adjacent methods

Further reading

Linked external resources are independent of TUCCA and Tufts University and remain under their own licenses.