Agrochemistry

Methodology of statistical analysis of biometric indicators of agricultural crops

For students

6 min read

Methodology of statistical analysis of biometric indicators of agricultural crops

To determine the linear correlation coefficients, a biometric analysis of plants is conducted based on traits, after which tables (matrices) of variation series for trait values are constructed.

To determine the correlation, other traits may also be used: plant height by developmental phases, leaf length, leaf width, leaf area, leaf area index, photosynthetic potential, accumulation of green (dry) mass of plants by developmental phases, and other quantitative plant traits, both biometric and biochemical. It is possible to determine correlation coefficients between the following characteristics: temperature regime, precipitation amount, growing season, nutrient content in the soil, number of productive stems, weeds, diseased plants, number of admixtures in cultivar crops; rates of seed sowing, application rates of mineral or organic fertilizers, pesticide application rates, seeding depth seeding depth seeding depth, crop density; and characteristics of tillage, as well as the physical, chemical, and water state of the soil. All of this can be measured, weighed, and calculated when computing correlation coefficients. 1026

Table 243 – Variation series for some rice traits

 No. Plant height, cm Number of productive stems Number of main panicles Length of panicle, cm Mass of grain from main panicle, g Total number of spikelets per main panicle Number of grains per main panicle Grain sterility (spikelets), % Mass of grain from main plant, g 1000 grain mass, g

1 2 3 4 5 6 7 8 9 10

1 69 2 13.5 2.68 109 92 2.54 15.59 4.42 27.60

2 67 3 12.8 2.26 108 74 1.97 31.48 5.36 26.62

3 79 4 15.3 3.75 151 133 3.58 11.92 10.65 26.91

4 67 2 14.3 2.88 110 98 2.71 10.90 3.53 27.65

5 67 2 15.1 2.93 113 102 2.77 9.73 4.29 27.15

6 75 3 14.2 3.22 123 107 3.02 13.00 6.92 28.22

7 76 3 14.7 3.28 130 117 3.09 10.00 8.18 26.41

8 66 2 13.0 2.12 90 78 2.00 13.33 3.28 25.64

9 78 3 15.7 3.90 153 135 3.69 1.76 9.40 27.33

10 65 3 14.4 2.75 128 93 2.49 27.34 5.41 26.77

Mandatory condition: the variation series of traits and the measurement sample must have the same quantity. If there are 10 of any characteristics in a matrix, then the number of measurements in each functional variable must be equal: 10, or 12, or 20 everywhere.

For determining intra-cultivar correlation links, one cultivar is sufficient. The number of plants (sample) can vary from 5 to 200 plants. With an increase in the plant sample size, the error of the correlation coefficient decreases and its significance (reliability) increases. To calculate inter-cultivar or genotypic correlation links, one needs to have trait values from several cultivars.

We determine linear correlation using the program or another program 6.0. For this, variation series of traits are entered from tables 242 and 243 one by one in order, starting from plant No. 1 and ending with plant No. 10. We obtain a correlation analysis matrix of the traits.

The PC produces similar matrices between all ten traits: 1–2; 2–3; 3–4; 4–5; 5–6; 6–7; 7–8; 8–9; 9–10.

In a report or a scientific article, one can use the values of correlation coefficients in the text. For example, the correlation coefficient between plant height and the number of spikelets in the panicle is r = 0.81±0.208.

The linear regression statistical model has the form: y = a+b·x, where the independent variable "x" is the argument or factor trait, and the dependent variable "y" is the function or resultant trait; regression coefficients "a" or "b" estimate the parameters of linear dependence in the entire general population. In the case of correlation of two traits, the variation of the resultant "y" values should be considered not only as a consequence of the dependence on the "factor" trait "x", but also as a result of random deviations.

Table 244 – Variation series for some rice traits

 No. Plant height, cm Number of productive stems Number of main panicles Length of panicle, cm Mass of grain from main panicle, g Total number of spikelets per main panicle Number of grains per main panicle Grain sterility (spikelets), % Mass of grain from main plant, g 1000 grain mass, g

1 2 3 4 5 6 7 8 9 10 1 73 3 15.7 3.52 126 110 3.31 12.69 8.93 30.09 2 75 4 17.0 4.37 163 145 4.15 11.04 13.77 28.62 3 67 2 14.5 3.27 139 117 3.0 15.82 4.09 25.64 4 77 4 16.1 3.32 136 103 2.93 24.26 9.38 28.44 5 71 3 16.0 3.68 127 11 3.39 12.59 9.18 30.54 6 61 2 14.7 2.34 104 80 2.17 23.07 3.80 27.12 7 70 3 14.5 2.86 163 98 2.51 38.65 6.23 25.61 8 72 4 15.0 2.97 146 103 2.062 29.45 8.43 26.11 9 69 3 15.7 3.02 120 102 2.90 15.00 6.18 28.43 10 77 4 15.8 3.11 107 95 2.93 11.21 9.28 30.84

To determine the value of any correlating trait in relation to another, a theoretical regression line is constructed on the correlation field. To construct a straight line on a graph, it is necessary to have two points. This statement is true for any two pairs of observations, regardless of whether the traits are linked to each other or not.

To construct the regression line, one should use the regression equations y = –3.33+0.08·73. To determine the significance (reliability) of the correlation coefficient, one should use the reliability values of the correlation coefficient.

Correlation coefficients for other cultivars are determined similarly.

The most common matrix, possessing full information capacity, is the matrix containing only direct correlation coefficient values.

Table 246 – Matrix of linear correlation coefficients between quantitative rice traits Correlating traits Traits 1 2 3 4 5 6 7 8 9 10

 2*) 0,88 3 0,67 0,64 4 0,60 0,43 0,79 5 0,23 0,41 -0,06 0,24 6 0,41 0,30 0,61 0,93 0,42 7 0,55 0,40 0,81 0,81 0,16 0,92 8 -0,22 -0,03 -0,61 -0,61 0,51 -0,43 -0,65 9 0,81 0,81 0,88 0,82 0,30 0,67 0,81 -0,36 10 0,51 0,35 0,71 0,43 -0,54 0,11 0,48 -0,72 0,54 1,00 *) Traits: see table 244.

To determine the reliability, one uses the values obtained when solving this problem, which are located in the matrices of tables 247 and 248.

We have been discussing phenotypic linear correlations. Y.L. Guzhov (1978) proposed using genotypic correlations. To construct a matrix of quantitative traits, one can use variation series of 5–7–10 cultivars. It is possible to take the mean trait values of 30–40–50 cultivars. From the mean trait values, variation series are constructed to calculate genotypic linear correlations. All other rules have already been described previously.

Table 247 – Matrix of correlation analysis results between rice traits Cor- Error Significance Regression values rela- Coeffi- of the Regression equation tion cient of correlation correlation (r) (m) (T) Type of a b c trait

 1–2*) 0,68 0,257 2,7 –3,33 0,08 0,38 у = а+в·х –3 0,61 0,279 2,2 6,58 0,11 0,64 –4 0,88 0,170 5,2 –3,61 0,09 0,14 –5 0,81 0,208 3,9 –86,37 2,93 7,28 –6 0,88 0,165 5,4 –135,35 3,36 4,76 –7 0,88 0,170 5,2 –3,85 0,09 0,14 –8 -0,47 0,312 1,5 62,03 -0,65 6,26 –9 0,92 0,135 6,9 –24,49 0,43 0,39 –10 0,24 0,343 0,7 24,75 0,03 0,73 *)

Traits: 1 – plant height; 2 – number of productive stems per plant; 3 – main panicle length; 4 – main panicle weight (grain+branches); 5 – number of spikelets in the main panicle; 6 – number of grains in the main panicle; 7 – grain weight per main panicle; 8 – grain sterility of the main panicle; 9 – grain weight per plant; 10 – 1000-grain weight.

Table 248 – Matrix of correlation analysis results between rice traits Cor- Coeffi- Error Significance Regression values rela- cient of of the Regression equation tion correla- correlation correlation (T) Type of a b c trait tion (r)

 1–2х) 0,88 0,165 5,4 –7,00 0,14 0,18 у = а+вх –3 0,67 0,262 2,6 7,53 0,11 0,47 –4 0,60 0,284 2,1 –1,49 0,07 0,37 –5 0,23 0,344 0,7 43,83 1,32 28,71 –6 0,41 0,322 1,3 3,67 1,44 14,85 –7 0,55 0,295 1,9 –1,36 0,06 0,40 –8 –0,22 0,345 0,6 49,24 –0,42 9,33 –9 0,81 0,210 3,8 –26,81 0,49 1,10 –10 0,51 0,305 1,7 13,51 0,21 0,51 *) Traits: see table 244.

When discussing the results of scientific research on linear phenotypic correlation, the coefficient of determination is often used – this is the square of the correlation coefficient. For example, in the Liman rice cultivar, the correlation coefficient between the traits of plant height and the number of grains per main panicle is r = 0,88, which is a close relationship between the traits. The square of the correlation coefficient is r2 = 0,882 = 0,77. This means that in 77% of cases, the number of grains in the panicle of the Liman rice cultivar is determined by the genotype, and in 23% of cases – by other conditions: stand density, plant leafiness, soil fertility, etc.

Multiple regression

To determine the linear phenotypic correlation, we conducted a sequential relationship between traits (example of 10 traits): 1–2; 1–3; 1–4; 1–5; 1–6; 1–7; 1–8; 1–9; 1–10; 2–3; 2–4; 2–5; 2–6; 2–7; 2–8; 2–9; 2–10; 3–4; 3–5; 3–6; 3–7; 3–8; 3–9; 3–10, etc.

In multiple regression, the presence of functional values is mandatory: x – any number; y – only one trait. For example, to determine the multiple regression for the Kuban 3 rice cultivar, we took the following x values: plant height (x1), main panicle length (x2), and number of spikelets per panicle (x3). For the y value, we used only one trait – the number of grains per main panicle (y). The sample size was 10 plants. Four variation series of 10 plants each were constructed. The results of the regression analysis were obtained:

Statistical processing of trait values (3 traits), we have – mean trait value, standard error of the mean, coefficient of variation, standard deviation, and experimental error;

Pairwise correlation coefficients between traits;

Multiple correlation coefficients;

Reliability of multiple correlation coefficients;

Contribution shares – x1y; x2y; x3y to the development of a specific trait.

Based on the statistical values of the regression analysis, a multiple regression equation is constructed, as well as a graph on which any point can be found depending on the actual values of x and y. Based on tables 250 and 251, we construct the multiple regression equation: у = а+в·х+с·х; у = 100,099+(– 2,803)·91,7+0,020·91,7. On the theoretical regression line, any point can be marked taking into account the x and y values, and its real value can be shown.

Table 249 – Results of statistical processing of traits

 Cor- Coeffi- Error Signif- Trait rela- cient of of icance x tion corre- the trait lation correlation 1. х1 91,7 1,30 4,5 4,1 1,4 х1-у 0,86 0,179 4,8 2. х2 18,2 0,36 6,2 1,1 2,0 х2-у 0,61 0,282 2,1 3. х3 102,0 3,21 9,9 10,1 3,1 х3-у 0,91 0,142 6,4 4. у 92,0 3,18 10,8 10,1 3,4 - - - -

Table 250 – Correlation relationship and path coefficients Pairwise correlation coef- Path coefficients Trait ficients between: x1 x2 x3 y x x1 x2 x3 influence share, %

 1. х1*) 1,00 0,38 0,67 0,86 х1 0,46 0,03 0,37 39,5 2. х2 0,38 1,00 0,62 0,61 х2 0,18 0,08 0,34 5,2 3. х3 0,67 0,62 1,00 0,91 х3 0,31 0,05 0,56 50,9 *) Traits: x1 – plant height; x2 – main panicle length; x3 – number of spikelets in the main panicle; y – number of grains in the main panicle

Table 251 – Results of regression analysis Statistical values Regression equation

 1) x1y; r = 0,81±0,207; T = 3,9 100,099; b = -2,803; с = 0,020 yx1; r = 0,82±0,199; T = 4,1 223,851; b = 1,617; c = 7889,8 2) x2y; r = 0,69±0,254; T = 2,7 117,542; b = -2,990; с = 0,021 yx2; r = 0,70±0,251; T = 2,8 227,784; b = 1,738; c = 8385,0 3) x3y; r = 0,89±0,100; T = 5,6 73,643; b = 0,579; с = 2505,5 yx3; r = 0,89±0,159; T = 5,9 2566,056; b = -71,153; с = 0,514

On the correlation field, one can plot the scatter of points (trait values at the intersection of coordinates) and the theoretical regression line based on the regression equation values.

Read next