Ridge/Lasso Lambda greater than 1
I ran Ridge and Lasso regressions using an algorithm to automatically find the optimum lambda. However, the algorithm couldn’t find an optimum lambda between 0 and 1. In some cases I could find optimum...
View ArticleCan L1 linear regression perform worse than vanilla linear regression on...
I have a data set with 2 features and I’m trying to predict one real-valued variable. I use linear regression and I measure the error using 10-fold CV and absolute mean error as a metric. I noticed...
View ArticleTikhonov regularization in the context of deconvolution
I came across “Tikhonov regularization” and I have bare knowledge on it. It seems that it is a type of regularization that is important for deconvolution. Are there any good resources and examples?...
View ArticleSelecting a loss-function for k-fold cross-validation over shrinkage parameter
I am doing a penalized regression with categorical (ordinal) outcomes. I would like to select the shrinkage parameter $lambda$ on the basis of cross-validation (CV). In this case, I have 50k...
View ArticleBayesian regularisation for ANNs — How to modify the Jacobian?
Introduction I have implemented the Levenberg-Marquard algorithm (from Hagan’s “Artifical Neural Network Design” — 2014) for a two layer network with 20 neurons in the hidden layer. This network can...
View ArticleDependence on UV cut off of some $phi^4$ diagrams
Consider the one loop corrections to the propagator and the vertex in $phi^4$-theory: The former gives an integral representation proportional to $int d^4...
View ArticleDifference between dropout and neurons with 0 weights
In the dropout method of regularization, we randomly delete half of the hidden neurons, leaving the input and output layers the same. In a theoretical sense, wouldn’t the same effect occur if we just...
View ArticleGMM via mclust in R builds models with only some cluster numbers,...
I am using the mclust package in R for Gaussian mixture modelling. On some data I encountered that the some model types were used only up to some cluster numbers. Here is a reproducible example:...
View ArticleWhy is it that xgb.cv performs well but xgb.train does not
I am trying to control overfitting using xgboost in R using eta but when I compare the overfitting of my xgb.cv readout to the xgb.train readout, I don’t know why xgb.cv doesn’t seem to overfit and...
View ArticleLasso with constraint on some coefficients (not all)
I would like to run a lasso regression (L1 penalisation) with a twist: there are different constraints on my problem. The coefficients for my features (predictors) are $beta_i$. I want to find the...
View ArticleIs my understanding of regularized logistic regression correct?
I learned that regularized logistic regression helps prevent the model from over-fitting the data. I understand that the function is still technically a high-order polynomial, but the effect is reduced...
View ArticleLASSO regression when model is known
I am very new to regression as I have been reading “The Elements of Statistical Learning: Data Mining, Inference, and Prediction” by Hastie et al. on Standford’s website this weekend. My goal is to...
View ArticleRepresenting propagators as Dirac delta functions [closed]
I have found online, in particular on the wolfram site, http://mathworld.wolfram.com/DeltaFunction.html, certain identities that allow one to represent a delta function as limits. Of particular...
View ArticleChoosing between feature selection and regularization to overcome...
In order to overcome over-fitting during a regression process over categorical features, one can either 1) Apply L1/L2/Elastic regularization during the regression, for example as answered here When to...
View ArticleRidge regression — why does the model only care to control large outliers?
One of the purposes of ridge regression is to curb the effects of outliers which may cause the regression coefficients to be so large and hence cause a highly biased model. That’s why the constraint...
View ArticleChoosing alpha for cost complexity pruning as described in Introduction to...
In the following lectures Tree Methods, they describe a tree algorithm for cost complexity pruning on page 21. It says we apply cost complexity pruning to the large tree in order to obtain a sequence...
View ArticleWhy do we need to normalize data before applying penalizing methods in the...
This question already has an answer here: Question about standardizing in ridge regression 1 answer
View ArticleWhy do smaller weights result in simpler models in regularization?
I completed Andrew Ng’s Machine Learning course around a year ago, and am now writing my High School Math exploration on the workings of Logistic Regression and techniques to optimize on performance....
View ArticleWhat is elastic net regularization, and how does it solve the drawbacks of...
Is elastic net regularization always preferred to Lasso & Ridge since it seems to solve the drawbacks of these methods? What is the intuition and what is the math behind elastic net?
View ArticleWhy is Lasso penalty equivalent to the double exponential (Laplace) prior?
I have read in a number of references that the Lasso estimate for the regression parameter vector $B$ is equivalent to the posterior mode of $B$ in which the prior distribution for each $B_i$ is a...
View Article