Wednesday, 31 August 2011

Discriminant Analysis Techniques and Segmentation


Discriminant Analysis techniques and Segmentation :
Market segmentation is one of the most essential factors for any product and brand. Market research is done prior before the launch. Segmentation variables have to be pre-selected and the data is collected, it is necessary to choose the statistical process by which the segments will be identified. The segmentation technique to be used depends largely on the type of data available (metric  or non metric variables), and the kinds of dependence observed - that is, dependence or  interdependence (Cooper D. & Emory, W., 1995, p. 521). Among the most common segmentation techniques used are factor analysis, cluster analysis, discriminant analysis, and multiple regression. Newer used techniques include chi-squared automatic detection (CHAID), LOGIT, and Log Linear Modeling  (Magidson, J., 1990).
The descriptive process is intended to yield a full bodied description of the market segments, which will be useful in the evaluation process but most importantly in the marketing mix creation stage. Multiple discriminant analysis is often used for this purpose (Gunter, B., & Furnham, A., 1992).
In marketing, discriminant analysis is often used to determine the factors which distinguish different types of customers and/or products on the basis of surveys or other forms of collected data. The use of discriminant analysis in marketing can be described by the following steps:
1.    Formulate the problem and gather data — Identify the salient attributes consumers use to evaluate products in this category — Use quantitative marketing research techniques (such as surveys) to collect data from a sample of potential customers concerning their ratings of all the product attributes.
    • The data collection stage is usually done by marketing research professionals. Survey questions ask the respondent to rate a product from one to five (or 1 to 7, or 1 to 10) on a range of attributes chosen by the researcher.
    • The attributes chosen will vary depending on the product being studied.
    • The data for multiple products is codified and input into a statistical program such as R, SPSS or SAS.
2.    Estimate the Discriminant Function Coefficients and determine the statistical significance and validity — Choose the appropriate discriminant analysis method.
    •  The direct method involves estimating the discriminant function so that all the predictors are assessed simultaneously.
    • The stepwise method enters the predictors sequentially.
    • The two-group method should be used when the dependent variable has two categories or states.
    • The multiple discriminant method is used when the dependent variable has three or more categorical states.
    • Use Wilks’s Lambda to test for significance in SPSS or F stat in SAS - The most common method used to test validity is to split the sample into an estimation or analysis sample, and a validation or holdout sample. The estimation sample is used in constructing the discriminant function. The validation sample is used to construct a classification matrix which contains the number of correctly classified and incorrectly classified cases. The percentage of correctly classified cases is called the hit ratio.
3.    Plot the results on a two dimensional map, define the dimensions, and interpret the results. The statistical program (or a related module) will map the results. The map will plot each product (usually in two dimensional space). The distance of products to each other indicate either how different they are. The dimensions must be labelled by the researcher. This requires subjective judgement and is often very challenging.

NEELIMA MAKANI
Marketing 2

Discriminant Function Analysis in the field of Supply Chain

Discriminant function analysis is used to determine which variables discriminate between two or more naturally occurring groups. For example, an educational researcher may want to investigate which variables discriminate between high school graduates who decide - to go to college, to attend a trade or professional school or to seek no further training or education For that purpose the researcher could collect data on numerous variables prior to students' graduation. After graduation, most students will naturally fall into one of the three categories. Discriminant Analysis could then be used to determine which variable(s) are the best predictors of students' subsequent educational choice. Suppose we have two groups of high school graduates: Those who choose to attend college after graduation and those who do not. We could have measured students' stated intention to continue on to college one year prior to graduation. If the means for the two groups (those who actually went to college and those who did not) are different, then we can say that intention to attend college as stated one year prior to graduation allows us to discriminate between those who are and are not college bound (and this information may be used by career counselors to provide the appropriate guidance to the respective students). Therefore, the basic idea underlying discriminant function analysis is to determine whether groups differ with regard to the mean of a variable, and then to use that variable to predict group membership
Such an analysis could also be used in the field of Supply Chain management. Supposing a company wants to determine the real impact of implementing Radio Frequency Indentification (RFID) in its business, it would have to investigate the various determinants on the adoption of such a technology. For example, the determinants can be taken as technological, organizational, environmental factors and product factors. The methodology adopted for a discriminant function would be to investigate the influential factors (independent variables) that may contribute to the RFID adoption (dependent variable). Discriminant analysis would be to determine whether statistically significant differences exist between the average score profile on a set of variables for prior defined groups and thereby enable those variables to be classified. It would also help to determine which of the independent variables account the most for the differences in the average score profiles of the two groups .  
If we code the two groups in the analysis as 1 and 2, and use that variable as the dependent variable in a multiple regression analysis, then we would get results that are analogous to those we would obtain via Discriminant Analysis. In general, in the two-group case we fit a linear equation of the type:
Y = a + b1*x1 + b2*x2 + ... + bm*xm
where a is a constant and b1 through bm are regression coefficients. The interpretation of the results of a two-group problem is straightforward and closely follows the logic of multiple regression: Those variables with the largest (standardized) regression coefficients are the ones that contribute most to the prediction of group membership.

Discriminant analysis-> a view point of an engineer

First I will start with the theory stuff which has become an important part of our life. Right from graduation to post graduation, theory is the way to go without even understanding its application in real world
So what Discriminant analysis is all about?
Discriminant analysis, or Discriminant function analysis, is a type of regression that focuses on membership in a naturally occurring group and the related predictor variable. It is a type of classification procedure that analyzes the individual characteristics of individual data points to determine which group the data points likely belong to. We must analyze a data set of known group membership before applying it to data set of unknown group membership.
We can use this tool in various situations such as
• How singles and married people differ?
• To predict which companies may go bankrupt based on financial variables. The classes being those who go bankrupt and those who don't.
• Facial recognition.
• Users and non users of a product.
But I have thought of two situations where we can also use them and it would be fun seeing the results.
Situation 1:
When today I worked on this tool I wondered it should have been taught to us after 12th standard at least to guys who wanted to pursue engineering. It could have been an important tool, as at that point of time we could use it to examine students who successfully complete their engineering degrees versus those who do not with different variables. Analysis could have been used to determine which variable, or combination of variables, predict whether a given student will successfully complete her engineering degree on time with two cases taken as students who complete on time and students who don’t.


Situation 2:
This analysis can also be applied to sports and being a true Indian the first sports that comes to my mind is cricket and the first sportsperson for whom the accuracy can be 100% is Sachin Tendulkar followed by Rahul Dravid. So by making two groups like Sachin and Rest of the Others we can apply this with variables such as technique, style, balance, statistic, composure and experience.


Name : Vivek Dhar
Group : Marketing 6

Discriminant Analysis

We use linear discriminant analysis when we classify objects into two or more groups based on knowledge of some variables or characteristics related to them. Discriminant analysis is somewhat similar to regression analysis. There is a dependant variable and some independent variables used to predict the independent variable in both the techniques. But in discriminant analysis, the dependant variable is categorical not metric.

Applications

1. Selecting MBA students based on some known scores in selection tests.

2. Purchase pattern of products in two categories – national brands and private labels. The independent variables taken for this could be annual income and household size.

3. Dividing a group of people into buyers and non – buyers

4. Selecting a potential candidate for the job or not.

5. It is used by credit rating agencies to rate individuals as high lending risk or low lending risk.

CALCULATION

Discriminant or discriminant function analysis is a parametric technique to determine which weightings of quantitative variables or predictors best discriminate between 2 or more than 2 groups of cases and do so better than chance (Cramer, 2003). The analysis creates a discriminant function which is a linear combination of the weightings and scores on these variables. The maximum number of functions is either the number of predictors or the number of groups minus one, whichever of these two values is the smaller.

Zjk = a + W1X1k + W2X2k + ... + WnXnk

Where:

Zjk = Discriminant Z score of discriminant function j for object k.

a = Intercept.

Wi = Discriminant coefficient for the Independent variable i.

Xj = Independent variable i for object k.

Again, caution must be taken to be clear that sometimes the focus of the analysis is not to predict but to explain the relationship, as such, equations are not normally written when the measures used are not objective measurements.

Name : Varun Aggarwal

Marketing Group 6

Discriminant Analysis

Discriminant Analysis is a Predictive method used to determine which continuous variables discriminate between two or more naturally occurring groups. Linear discriminant analysis divide the variables into two groups. For example a researcher may want to find out which people would likely to yes to some particular question and which people would likely to say no to that particular question. Discriminant analysis thus helps to predict the responses of people to that particular question that have missed to answer it. The prediction is done on the basis of various parameters on which the response would be dependent on. Those parameters will be called as independent variables while the answer to the question will be called as dependent variable.

Thus discriminant analysis is a regression analysis when only 2 groups are involved. A transfer function is formed on the basis of dependent and independent variables as shown below:

Y = aX1 + bX2 + c X3 + dX4 + …….. +nXm

Where: Y = output score

X1, X2,X3…Xm = Independent variables

a, b, c ….n = co-efficients of Xs which are actually the weights of various independent variables (or predictor variables)

These co-efficients are also called as standardized co-efficients. However, these coefficients do not tell us between which of the groups the respective functions discriminate. We can identify the nature of the discrimination for each discriminant function by looking at the means for the functions across groups.

Data of this type may be represented in any number of different forms: scatterplots, tables of means and standard deviations, and overlapping frequency polygons.

As previously mentioned, DA is usually used to predict membership in naturally occurring groups. It answers the question: can a combination of variables be used to predict group membership? Usually, several variables are included in a study to see which ones contribute to the discrimination between groups. The means for the significant discriminant functions are examined in order to determine between which groups the respective functions seem to discriminate.

DA can be used for Research and Development, Consumer and market research, quality control and quality assurance across a range of industries such as food and beverage, paint, pharmaceuticals, chemicals, energy, telecommunications, and others. It can generate data models not only for prediction but for faster product and process optimization for application in Product Development and Quality Control. Apart from this, DA is used in speech recognition software.

DA can be extensively used in the Six Sigma projects. Inside the Improve phase of Six Sigma, while identifying and testing various solutions we need to develop transfer function to establish a relationship between the project output variables and various input variables. Once we plot the variables on scatter diagrams, pareto charts we could be able to draw the cause and effect diagrams identifying the various Xs and Ys. Once we verify the various root causes in the Analyse phase of the six sigma project, we could conduct design of experiments (DoE) and regression analysis to predict the output (Y).

Author - Siddhartha Sabale

Group 2 _ Operations

Of faulty assumptions and more

Rationality and normal distribution of factors are the two assumptions that underlie most theories in finance. The most frequently quoted but unrealistic, may I add!!

So I came across this technique that uses discriminant analysis to predict the failure of a bank. And if the two assumptions mentioned above were actually true, we would have beaten the financial crisis hands down and saved the financial giants from going down in the dumps. Having said that though, this technique is at least worth discussing.

Firms are divided into failing and non-failing firms and each is given a ‘discriminant score’. The technique works as follows:
Discriminant analysis is used to derive a linear combination of two or more independent variables that best discriminate the two groups under consideration, failing and non-failing companies. The discriminant analysis derives a linear combination from the following equation

Z = w1x1+ w2x2+...+wnxn
Where

Z = discriminant score
wi (i=1, 2, ... ,n) = discriminant weights
xi (i=1, 2, ... ,n ) = independent variables, that are various financial ratios in this case

The discriminant score given to each firm is then compared to a cut-off score to determine which group the company in question belongs to (failing or non- failing firm).

Now, in the perfect world where everything followed the ‘norm’, discriminant analysis would have worked perfectly given that the variables in every group followed a normal distribution and the covariance matrices for every group were equal. However, past experiments have revealed that firms and more often than not, failing firms, often violate this condition. So an aberration from the assumption is more of a norm than just an exception.

However, being the ‘bean counters’ that we finance guys are accused of being, a favourable discriminant score should make for a good enough premise for us to declare a firm failing or non- failing. No room for faulty assumptions there!

Reference : Paper titled ‘Choosing Bankruptcy predictors,using Discriminant Analysis, Logit Analysis and Genetic Algorithms’ by the Turku Centre for Computer Science

Posted by
Jyoti Maheshwari
Finance

Using discriminant analysis to prevent bankruptcy

Earlier Ratio Analysis was used as an analytical technique to measure the performance of the companies. The performance of the companies, as measured by the ratios , such as profitability, liquidity and solvency prevailed as the most significant indicator but the order of their importance is not being clear , since in each and every case taken ,some different ratio came out as the most important factor and sometimes it might seem to be confusing , for instance, a firm with a poor profitability / solvency record may be regarded as potentially bankrupt.

Discriminant Analysis is a technique , which has been applied nowadays successfully to financial problems such as credit evaluation (as had been illustrated in class) and investment analysis. It is a parametric technique to determine , which weightings of quantitative variables or predictors best discriminate between 2 or more than 2 groups of cases. Analysis creates a discriminant function , which is a linear combination of the weightings and scores on these variables. The maximum no. Of functions is either the no. Of predictors or the no. of groups minus one , whichever of these two values is the smaller.

Z =a+W1*X1+W2*X2+.......+Wn*Xn

Where Zjk = Z score of discriminate function

a= intercept

Wi=Discriminant coefficient for the independent variable i

Xi =independent variable i

In a 2 group discriminant function, cutting score is used to classify he two groups uniquely. Optimal cutting score is halfway between the centroid of two groups.

The primary goal is to find dimensions that groups differ on and ,create classification functions , then we need to find along how many dimensions do groups differ reliably.

By discriminant analysis , we can find which ratios are affecting the performance of a company .

Z= a*X1+b*X2+c*X3+....+n*Xn

Where Xis can be working capital / total assets, retained earnings / total assets, EBIT/total assets, B/E value, sales/ total assets etc.

It can then be applied to similar other companies.

We use discriminant analysis in evaluating consumer loan application. A fast and efficient device for detecting unfavourable credit risks might avoid the loan officer to avoid disastrous decisions. Thus it can be thought of as a potential weapon to be used in business sector.


Posted by

Neeraj Kumar Singh

Finance