Scripts for general data analyses (Laboratory purposes) and some algorithms for learning purposes
(Principal Component Analysis python scripts)
Assume the presence of an arbitrary 2D Gaussian distribution within a two-dimensional space.
Let us sample 1,000 random points from this 2D Gaussian distribution.
Consider each point as a set of vectors in a two-dimensional space:
The covariance matrix for
, where
Define a projection matrix for the linear transformation
The variance of the projected points onto
We aim to maximize the variance such that the projected points correspond to the principal component.
To prevent the variance from diverging, we fix the norm of the projection matrix at 1. Therefore, we solve the following constrained maximization problem to find the projection axis:
To solve this maximization problem under the given constraint, we employ the method of Lagrange multipliers. This approach introduces an auxiliary function, known as the Lagrange function, to find the extremum of a function subject to constraints. The formulation of the Lagrange multipliers method applied to this problem is:
At the point of maximum variance:
Hence, we find that:
Select the eigenvector corresponding to the largest eigenvalue (where
A singular value decompositon of
, where
Right Singular Vectors:
thus
Right Singular Vectors:
thus
Hence,
Then, we can derive the following:
We can solve this equation as follows:
Therefore, the singular value is described as:
where
Assume that we have two series of data that describe expression patterns of stress response genes.
data1:
| rpoS | dnaK | oxyR | sosR | cspA | |
|---|---|---|---|---|---|
| 0 | 0.61 | 0.16 | 0.08 | 0.05 | 0.10 |
| 1 | 0.54 | 0.18 | 0.07 | 0.10 | 0.12 |
| 2 | 0.57 | 0.10 | 0.06 | 0.15 | 0.12 |
| 3 | 0.39 | 0.16 | 0.17 | 0.13 | 0.15 |
| 4 | 0.45 | 0.14 | 0.18 | 0.14 | 0.08 |
| 5 | 0.42 | 0.08 | 0.12 | 0.19 | 0.19 |
| 6 | 0.50 | 0.17 | 0.05 | 0.16 | 0.11 |
| 7 | 0.44 | 0.16 | 0.08 | 0.18 | 0.15 |
| 8 | 0.51 | 0.12 | 0.23 | 0.05 | 0.09 |
| 9 | 0.59 | 0.04 | 0.12 | 0.05 | 0.20 |
data2:
| rpoS | dnaK | oxyR | sosR | cspA | |
|---|---|---|---|---|---|
| 0 | 0.19 | 0.60 | 0.01 | 0.08 | 0.12 |
| 1 | 0.19 | 0.44 | 0.12 | 0.15 | 0.11 |
| 2 | 0.14 | 0.46 | 0.18 | 0.11 | 0.11 |
| 3 | 0.11 | 0.61 | 0.08 | 0.14 | 0.07 |
| 4 | 0.05 | 0.53 | 0.15 | 0.16 | 0.12 |
| 5 | 0.02 | 0.50 | 0.16 | 0.16 | 0.16 |
| 6 | 0.16 | 0.54 | 0.10 | 0.07 | 0.13 |
| 7 | 0.04 | 0.53 | 0.16 | 0.17 | 0.10 |
| 8 | 0.17 | 0.46 | 0.08 | 0.12 | 0.17 |
| 9 | 0.19 | 0.57 | 0.02 | 0.12 | 0.10 |
Using the raw data, it is impossible to visualize as the data dimension is higher than 3. Thus, try to reduce the dimensionality.
As the first step, let us visualize the singular values for each dimension.
Since the cumulative contribution rate exceeds 90% up to three dimensions, we will project the data into a three-dimensional space for analysis.
Upon projecting the data into a three-dimensional space, it becomes evident that the gene expression patterns of data1 and data2 are distinctively separated.
The specific growth rate is a coefficient that represents how much one cell divides per unit of time.
Under constant conditions, with the cell mass denoted as X (g), the specific growth rate as μ (
thus
where
Here, we calculate the specific growth rate (
Here is an example of sequential OD600 data of an Escherichia coli strain for 20 hours.
The exponential phase in this case is where OD600 is in between 0.01 and 0.10, therefore we only focus on the area.
Assume that the cell math icreases exponentially at the phase, we set a fitting model given by:
is also written as:
thus we use a normal equation:
where
With the model, we obtained the result where the specific growth rate is:
and the fitted growth curve is as shown below.
In this section, we calculate Oligonucleotide Melting Temperature, denoted as Tm, with Nearest Neighbors method.
With Nearest Neighbors method, Tm is calculated as:
, where
Tm = Melting temperature in ℃
A = constant of -0.0108
R = gas constant of 0.00199
C = oligonucleotide concentration in
and the parameters are already emprically determined as shown:
| AA | -9.1 | -0.024 |
| AT | -8.6 | -0.0239 |
| TA | -6.0 | -0.0169 |
| CA | -5.8 | -0.0129 |
| GT | -6.5 | -0.0173 |
| CT | -7.8 | -0.0208 |
| GA | -5.6 | -0.0135 |
| CG | -11.9 | -0.0278 |
| GC | -11.1 | -0.0267 |
| GG | -11.0 | -0.0266 |
| TT | -9.1 | -0.024 |
| TG | -5.8 | -0.0129 |
| AC | -6.5 | -0.0173 |
| AG | -7.8 | -0.0208 |
| TC | -5.6 | -0.0135 |
| CC | -11.0 | -0.0266 |
- Create a virtual environment for python3
python3 -m venv venv- Activate the venv
source venv/bin/activate- Leave the environment
deactivatepip freeze > requirements.txtpip install -r requirements.txt






