Application of Matrices, Regression, and ANOVA in the Design and Analysis of Experimental Data Using Statistical Software for Pharmaceutical Formulation Optimization

Authors:
  • Chaithra N , Division of Medical Statistics, School of Life Sciences, JSS Academy of Higher Education & Research, Mysuru, Karnataka, India. ORCID id: 0000-0001-5824-9800
  • Panchaksharappa Gowda D.H , Department of Pharmacy Practice, JSS College of Pharmacy, JSS Academy of Higher Education & Research, Mysuru, Karnataka, India. ORCID id: 0000-0003-4282-3040
  • Chethan I.A , Department of Pharmaceutical Chemistry, JSS College of Pharmacy, Mysuru, Karnataka, India. ORCID id: 0000-0002-0589-1362
  • Anil Kumar K M , Department of EnvironmentalScience, School of Life Sciences, JSS Academy of Higher Education & Research, Mysuru, Karnataka, India.

Article Information:

Published:December 26, 2025
Article Type:Original Research
Pages:2288 - 2299
Received:September 28, 2025
Accepted:November 24, 2025

Abstract:

In the early twentieth century, we were analyzing the experimental data after performing the experiment, but in the meantime the idea of using statistical analysis in the planning stages of research, as opposed to the conclusion of experimentation, was first presented by Sir Ronald Fisher. This concept of planning, designing, performing and analyzing the experimental results was used in the field of Engineering and Agriculture Science. In over time the pharmaceutical industries started applying these new paradigms extensively in the formulation development to optimize the product, to measure the main effect and interaction effect of factors or ingredients selected in preparing a formulation and to verify the level of significance between the experimental values which are obtained after performing and predicted values which are computed after obtaining the mathematical model or mathematical equation. The objective of this study was to help the investigator understand the application of matrix techniques, Standard Deviation, Coefficient of Variance, Regression and ANOVA for analyzing the data of different formulations prepared using different ingredients at different quantity or levels. To perform these types of studies in pharmaceutical research, the Design of Experiments (DoE) is a main statistical tool, that can be used for analyzing the data and to verify the level of significance α = 0.05 or 0.01 by determining F-ratio. This statistical analysis shows that the model which is developed is more appropriate to perform experiments and analyze the resulting data.

Keywords:

Design of Experiments Application of Matrices Regression and ANOVA.

Article :

INTRODUCTION:

The primary goal is to demonstrate how Mathematical and Statistical methods are more appropriate to analyze the response variable using DOE (Design of Experiments) [1]. It can be applied for Matrices & Determinants, Standard Deviation, Coefficient of Variation, Regression and ANOVA. To perform this type of study, DOE is the only statistical software which can be used to Design, Perform the experiment and analyzing the result to verify the level of significance [2]. Before the existence or development of DOE, investigators were using the traditional method of changing one factor at time among the input variable within an appropriate levels by retaining the range of remaining factors as constant. But the traditional method OFAT (Integrated Factory Acceptance Test) approach fails to explain the measure of interaction between the   factors, which may lead to an inadequate conduction of the development and optimization. To overcome these types of problems, DoE has been considered as a better approach to get accurate results with few numbers of experiments [3]. DoE has become one of the main components of pharmaceutical and analytical QbD. In this paper we are providing the theoretical and practical aspect of DoE in pharmaceuticals [4]. There are different software’s such as GraphPad Prism, SPSS, R, and MATLAB are utilized for conducting statistical analyses These software packages are often found to be costly, and their utility extends beyond statistical analyses to encompass curve fitting, modeling, and graphing functionalities, which are more expensive if the user's only objective is to perform the statistical analysis [5].

MATERIAL AND METHODS:

Computational methods and theory

To prepare pharmaceutical formulation, investigators may consider n-ingredients as input variables and they perform an experiment to obtain one response variable or formulation. In general, the input variables which are selected for study will be designated as low and high level and the designations of the variables will be noted as -1 for low level and +1 for high level of any one variable. But for the response variables, one cannot assign the coded values, here one has to consider actual value of the resultant [6]. If the investigator considers the three input variables, namely A, B and C, their designated levels can be tabulated as.

 

Table 1: Coded Levels of Ingredients A, B, and C

INGREDIENTS

DESIGNATED LEVELS IN CODED FORM

Low level

High Level

A

-1

+1

B

-1

+1

C

-1

+1

If the investigator considers the factorial design in DOE to design and analyze the data of a formulation, then 23 – Factorial design has to be chosen to perform experiments [7].

 

Table 2: Design of experiment without interaction and response variable.

Runs

Combinations

of ingredients with different levels

Ingredients

A

B

C

1

1

-1

-1

-1

 

2

a

1

-1

-1

 

3

b

-1

1

-1

 

4

ab

1

1

-1

 

5

c

-1

-1

-1

 

6

ac

1

-1

1

 

7

bc

-1

1

1

 

8

abc

1

1

1

 

 

 

Mathematical and statistical methods to obtain a Model to analyze the data in DOE

 

To analyze the experimental results, investigators should follow the matrix method to obtain the mathematical or regression model. To obtain regression model, one has to proceed as per the steps listed [8].

 

Step 1: Construct three matrices as per ingredients or input variables selected, response variables and regression coefficients need to be computed.

 

X =     =       Y =      =

Table 3: Design of experiment table with response variable.

 

Runs

Combinations

of ingredients with different levels

Ingredients

A

B

C

Response Variable – Y1

1

1

-1

-1

-1

Y1

 

2

a

1

-1

-1

Y2

 

3

b

-1

1

-1

Y3

 

4

ab

1

1

-1

Y4

 

5

c

-1

-1

-1

Y5

 

6

ac

1

-1

1

Y6

 

7

bc

-1

1

1

Y7

 

8

abc

1

1

1

Y8

 

 

RESULTS AND DISCUSSION:

Step 2: Obtaining the transpose Matrix of X

 

XT =

 

 

Step 3: Obtaining the product of  XTX

 

XTX = *      

 

Model example of multiplication of matrices

                                   …..(1)

          

                

                                   …..(2)

          

                                    

 

XTX =  

XTX =  

Step 4: Obtaining the co-factor matrix of XTX

 

A =

Let                                                                                       

Co-factors are

    ,

, ,

, ,

      The  C.F matrix A =     =

Step 5: Obtaining the adjoint of XTX matrix as explained in the

      The  C.F matrix A =     =

 

adj

Step 6: Obtaining the value of determinant A

               According to the property of determinant

a11A11 + a12A12 + a13A13 = |A|

 =

Step 7: Obtaining the inverse of XTX =(XTX)-1

 

              A = 

               =

The  C.F matrix A =     =

a11A11 + a12A12 + a13A13 = |A|

      1*3 + 4*6 + (-2)*9  = 9

In the given example A-1 = adjA  =  

Step 8: Obtaining the inverse of XTY

Step 9: Obtaining the product of (XTX)-1XTY =  =

Then the regression equation can be written as Y1 =

Step 10: Obtaining the sum of the squares due to regression, due to error and Total sum square

Runs

Y

Y1=

(Y - Y1)2

(Y1 -  )2

(Y - 2

1

Y1

(Y1 -  )2

(   -  )2

(Y1 - 2

2

Y2

(Y2 -  )2

(   -  )2

(Y2 - 2

3

Y3

(Y3 -  )2

(   -  )2

(Y3 - 2

4

Y4

(Y4 -  )2

(   -  )2

(Y4 - 2

5

Y5

(Y5 -  )2

(   -  )2

(Y5 - 2

6

Y6

(Y6 -  )2

(   -  )2

(Y6 - 2

7

Y7

(Y7 -  )2

(   -  )2

(Y7 - 2

8

Y8

(Y8 -  )2

(   -  )2

(Y8 - 2

 

 

 

= Sum of squares due to residual error

SSE

= Sum of squares due to regression

SSR

 

= Total sum squares

SST

   Y1 + Y2 + Y3 + Y4 + Y5 + Y6 + Y7 + Y8

 =

N= Number of runs

Step 11: Developing of ANOVA table

Table 4: Construction of ANOVA table to verify the level of significance of Regression

Sources of variation

Sum of squares

Degree of freedom

Mean squares

F – Distribution value (F-Ratio)

Variation due to regression

SSR = Sum of squares due to regression

λ1 = 3 (Degree of freedom is equal to number of independent variables chosen for the study)

MSR = Mean square due to regression

MSR =

 

F=

Variation due to error

SSE = Sum of squares due to error

λ2 = n- 1+3

λ2 = 8 1+ 3

λ2 = 4

MSE = Mean square due to error

 

Total

SST

λ3 = n - 1

= 8 - 1 = 7

 

 

Step 12: Analyzing the main and interaction effect of factors:

Investigator has to verify main effect and interaction effect of the ingredients which are chosen for preparing the formulation [9]. The regression equation which is obtained by applying the matrix method is taken as Y1 = , here  ,  and   are regression coefficients of input variables. Here we have taken the regression equation without an interaction [10].

If the interaction term is statistically significant and p<0.05 the interaction term has to be considered to  analyze the response variables and also if the coefficient interaction term is very high, then the investigator has to consider the interaction effect and the model will be  Y1 = If P > 0.05, then the interaction effect has to be removed from the model equation [11,12].

 

 

 

 

 

Table 5: Combinations of ingredients with interaction

 

Runs

Combinations

of ingredients with different levels

Ingredients with interaction

 

A

B

C

AB

AC

BC

ABC

Response Variable is Y1

1

1

-1

-1

-1

1

1

1

-1

Y1

2

a

1

-1

-1

-1

-1

1

1

Y2

3

b

-1

1

-1

-1

1

-1

1

Y3

4

ab

1

1

-1

1

-1

-1

-1

Y4

5

c

-1

-1

1

1

-1

-1

1

Y5

6

ac

1

-1

1

-1

1

-1

-1

Y6

7

bc

-1

1

1

-1

-1

1

-1

Y7

8

abc

1

1

1

1

1

1

1

Y8

The main (Average) effect of each of the factors can be obtained using the mathematical equations

A =  = 

B =  

C=  =

 here n= number taken as replication.

The interaction effects can be obtained as

AB =  =

BC =  =

AC =  =

ABC =  =

Step 13: The contrast is an important part of the experimental design. It describes the key differences that are being compared within an experiment.

The contrast of the first factor or variable A is obtained using the equation

Usually, this contrast in DOE is called as total effect. Similarly, investigators can obtain the contrast for factor B, C, AB, AC, BC and ABC.

The sum of the squares for each effect are computed, because each effect will have a single degree of freedom contrast [13 -16].

The sum of squares of the factor A =  

                                                           =

The sum of squares of the factor B=  

                                                           =

The sum of squares of the factor C=  

                                                           =

The sum of squares of the factor AB=  

                                                           =

The sum of squares of the factor BC=  

                                                           =

The sum of squares of the factor AC=  

                                                           =

The sum of squares of the factor ABC=  

                                                           =

Table 6: 23 - Factorial Design

 

 

Factor 1

Factor 2

Factor 3

Response 1

Std

Run

A: Stearate

B: Drug

C: Starc

Thickness (Mg)

4

1

1.5

120

30

420

5

2

0.5

60

50

500

1

3

0.5

60

30

350

3

4

0.5

120

30

210

6

5

1.5

60

50

530

2

6

1.5

60

30

428

8

7

1.5

120

50

505

7

8

0.5

120

50

455

Table 7: The ANOVA table for the Statistical model of thickness of the response variable.

 

Source

Sum of Squares

df

Mean Square

F-value

p-value

Model

65209.00

3

21736.33

8.18

0.0350

A-Stearate

16928.00

1

16928.00

6.37

0.0650

B-Drug

5940.50

1

5940.50

2.24

0.2091

C-Starc

42340.50

1

42340.50

15.94

0.0162

Residual

10624.50

4

2656.13

 

 

Cor Total

75833.50

7

 

 

 

The model is significant, as indicated by the F-value of 8.18. The probability of an F-value of this magnitude by random chance is only 3.50%. Model terms are considered significant when their p-values are less than 0.0500. In this specific instance, "C" emerges as a significant model term, while "B" does not achieve significance, as its p-value exceeds 0.1000.

 

 

Table 8: Fit Statistics

Standard Deviation

51.54

Mean

424.75

C.V.% - Coefficient of Variation

12.13

R2

0.8599

The coefficient of variation is obtained as C.V.% =  

                                                                           = (51.54/424.75) = 12.13%

The coefficient of variation is 12.13%, indicating the degree of variation of the of response variables of  a particular series. If the investigators observe, the coefficient of variation is very high, indicates that the design of experiment need to be modified or not appropriate [17,18].

 

After doing the analysis, model was found significant. Final Equation in Terms of Actual Factors. Y = 123. 50 + 92.00X1 – 0.908X2 +7.2750X3.

 

The equation, expressed in terms of actual factors, enables predictions about the response for given levels of each factor. In this case, each factor levels ought to be stated in their original units. However, this equation should not be utilized to assess the relative impact of each factor due to the coefficients being scaled to accommodate the units of each factor, and the intercept not being positioned at the center of the design space [19 - 23].

 

Report of Actual and Predicted value are obtained using the model equation Y = 123. 50 + 92.00X1 – 0.908X2 +7.2750X3.

 

Table 9: Regression Model Residuals

Run Order

Actual Value

Predicted Value

Residual

1

420.00

370.75

49.25

2

500.00

478.75

21.25

3

350.00

333.25

16.75

4

210.00

278.75

-68.75

5

530.00

570.75

-40.75

6

428.00

425.25

2.75

7

505.00

516.25

-11.25

8

455.00

424.25

30.75

 

 

Figure 1:  Graph to how observed values scattered around the predicted values (around the straight line) in DOE.

 

Figure 2:  Graph to how observed values scattered around the predicted values (around the straight line) in JMP software.

 

 The graph which is constructed taking the observed values on X-axis and Predicted Values on Y-Axis indicated the observed /experimental values are very close to the straight line of predicted values. This is an indication to show that model is significant [24]. Distributions of response variables is normal.

 

Figure 3:  Graph to show the normality of the distribution.

 

CONCLUSION:

The following conclusion was based on the above study was to explain the important applications of Matrices in developing Regression model or forming the regression equation in establishing the functional the relationship between the input variables (Ingredients which are used with different levels to prepare a formulation) & response variable (Output Variable). After developing the regression model, it is explained how we can verify the effect of each variable, the interaction effects of between the variables or ingredients and explained the importance of ANOVA to verify the level of significance of the mathematical model.

 

Competing Interests

The authors declare no competing interests

 

Authors’ Contributions

All authors made contributions to the writing of this article and have approved the final manuscript

 

REFERENCES:

1.     Jankovic A, Chaudhary G, Goia F. Designing the design of experiments (DOE) – An investigation on the influence of different factorial designs on the characterization of complex systems. Energy Build. 2021;250(111298):1-7.

2.     Ranga S, Jaimini M, Sanjay K, Sharma B, Singh A. A review on design OF experiments (DOE) 2014;3(1):216-224

3.     Choguill CL. The research design matrix: A tool for development planning research   studies. Habitat Int. 2005;29(4):615–26.

4.     Koukouvinos C, Seberry J. Weighing matrices and their applications. J Stat Plan Inference.  1997;62(1):91–101.

5.     Yu P, Low MY, Zhou W. Design of experiments and regression modelling in food flavour and sensory analysis: A review. Trends Food Sci Technol. 2018; 71:202–15.

6.     Khomchenko AN. Applied topics of control theory, mathematical statistics, and mathematical cybernetics: A version of the projective grid method. J Math Sci. 1993;66(5):2525–7.

7.     Tovmachenko NN, Fedorov VV. Regression analysis in experimental design problems. J  Math Sci. 1992;60(4):1607–10.

8.     Rodriguez-Granrose D, Jones A, Loftus H, Tandeski T, Heaton W, Foley KT, et al. Design of experiment (DOE) applied to artificial neural network architecture enables rapid bioprocess improvement. Bioprocess Biosyst Eng. 2021;44(6):1301–8

9.     Schubert J, Simutis R, Dors M, Havlik I, Lübbert A. Bioprocess optimization and control: Application of hybrid modelling. J Biotechnol. 1994;35(1):51–68.

A.     Sethuramiah, Rajesh Kumar. Chapter 6 - Statistics and Experimental Design in Perspective.Modeling of Chemical Wear.Elsevier.2016 Pages 1-31,

 

10.   Jankovic A, Chaudhary G, Goia F. Designing the design of experiments (DOE) – An investigation on the influence of different factorial designs on the characterization of complex systems. Energy Build. 2021;250(111298):111298

11.   Fan Y, Chen J, Shirkey G, John R, Wu SR, Park H, et al. Applications of structural equation modeling (SEM) in ecological studies: an updated review. Ecol Process. 2016;5(1)

12.   Fukuda IM, Pinto CFF, Moreira C dos S, Saviano AM, Lourenço FR. Design of experiments (DoE) applied to pharmaceutical and analytical quality by design (QbD). Braz J Pharm Sci. 2018;54:01006

13.   Brown AM. A new software for carrying out one-way ANOVA post hoc tests. Comput Methods Programs Biomed. 2005 Jul;79(1):1-7.

14.   J. D. Faucher and P. Maussion, "Response Surface Methodology for the Tuning of Fuzzy Controller Dedicated to Boost Rectifier with Power Factor Correction," 2006 IEEE International Symposium on Industrial Electronics, Montreal, QC, Canada, 2006, pages-1-6.

15.   Conte N, Díez E, Almendras B, Gómez JM, Rodríguez A. Sustainable recovery of cobalt from aqueous solutions using an optimized mesoporous carbon. J Sustain Met. 2023;9(1):266–79.

16.   Porto de Lima B, da Silva AF, Marins FAS. New hybrid AHP-QFD-PROMETHEE decision-making support method in the hesitant fuzzy environment: an application in packaging design selection. J Intell Fuzzy Syst. 2022;42(4):2881–97.

17.   Sartori MMP, Florentino H de O. Energy balance optimization of sugarcane crop residual biomass. Energy. 2007;32(9):1745–8.

18.   Freij A, Yuan S, Zhou H, Solihin Y. Persist level parallelism: Streamlining integrity tree updates for secure persistent memory. In: 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture. IEEE; 2020. p. 14–27.

19.   Jauregui M, Vílchez J, Chacón L. A PROCEDURE FOR MAP UPDATING USING DIGITAL MONO-PLOTTING AND DTMs. Isprs.org.

20.   Wazwaz AM. A new method for solving singular initial value problems in the second-order ordinary differential equations. Applied Mathematics and computation. 2002;10;128(1):1-13.

21.   Ghellam M, Koca İ. Autumn olive berries: characterization, antioxidant capacity, optimization of solvents and extraction conditions of lycopene. Journal of Food Measurement and Characterization. 2021:1-11.

22.   Rouberty F, Fournier J. Modelling of GC and HPLC separations of simazin and atrazin by experimental design methodology. Chromatographia. 1995;41(9-10):1-8.

23.   Beane G, Geuther BQ, Sproule TJ, Trapszo J, Hession L, Kohar V, Kumar V. Video based phenotyping platform for the laboratory mouse. bioRxiv. 2022; 18:20 22-01.

24.   Yu P, Low MY, Zhou W. Development of a partial least squares-artificial neural network (PLS-ANN) hybrid model for the prediction of consumer liking scores of ready-to-drink green tea beverages. Food Research International. 2018;103:68 - 75.

25.   Vats A, Sharma VS, Ansari I. Effects of parameters on burr heights and diametral error in dry drilling. International Journal of Additive and Subtractive Materials Manufacturing. 2017;1(3-4):223-39.

26.   Roy C, Chakrabarty J. Quality by design-based development of a stability-indicating RP-HPLC method for the simultaneous determination of methylparaben, propylparaben, diethylamino hydroxybenzoyl hexyl benzoate, and octinoxate in topical pharmaceutical formulation. Scientia Pharmaceutica. 2014;82(3):519-40.

27.   Rao Lakkimsetty N, Karunya S, Kavitha G, Ali Moosa Al Balushi N, Sulaiman Ali Al Rashdi S, Shaik F. RETRACTED: Application of waste tire derived activated carbon for the removal of heavy metals in wastewater using Response surface methodology. Mater Today. 2023.

28.   Bakri MK, Nyuk Khui PL, Rahman MR, Hamdan S, Jayamani E, Kakar A. Dielectric Properties of Acacia Wood Bio-composites. Acacia Wood Bio-composites: Towards Bio-Sustainability of the Environment. 2019:171-86.

29.   Vats A, Sharma VS, Ansari I. Effects of parameters on burr heights and diametral error in dry drilling. International Journal of Additive and Subtractive Materials Manufacturing. 2017;1(3-4):223-39.

30.   Kacan E. Exergetic optimization of basic system components for maximizing exergetic efficiency of solar combisystems by using response surface methodology. Energy and Buildings. 2015;15;91:65-82.

 

31.   Zhao L, Chen Y, Schaffner DW. Comparison of logistic regression and linear regression in modeling percentage data. Applied and environmental microbiology. 2001;1;67(5):2129-35.