From: keith steward
[ksteward@bacterialbarcodes.com]
Sent: Wednesday, December 17, 2003
5:14 PM
To: Home
Subject: {Blocked Content} FW: rep-PCR data
analysis
review and do background reading as needed to better
comment
-----Original Message-----
From:
David.Renter@gov.ab.ca [mailto:David.Renter@gov.ab.ca]
Sent:
Wednesday, December 17, 2003 3:14 PM
To: Alex Renwick;
John.Berezowski@gov.ab.ca
Cc: Mimi Healy; jmanry@bacbarcodes.com;
arenwick@bacbarcodes.com; ksteward@bacbarcodes.com
Subject: Re:
rep-PCR data analysis
Alex,
Thank you for your detailed
reply. It sounds like you understand our different approach and have
thought about some of these issues. John and I have provided some
clarifications and further questions below. (- John please add to the
discussion if I missed anything) I'm not sure we have reached a solution, but I
believe we are on the right track.
Thanks again
Dave
David G. Renter, DVM, PhD
Veterinary
Epidemiologist
Agri-Food Systems Branch
Food Safety Division
Alberta
Agriculture, Food and Rural Development
Phone: 780/422-0275
Fax:
780/422-3438
|
| "Alex Renwick"
<arenwick@bacbarcodes.com>
12/16/2003 11:35 AM
| To:
David.Renter@gov.ab.ca cc: "Mimi
Healy" <mhealy@bacbarcodes.com>,
jmanry@bacbarcodes.com;arenwick@bacbarcodes.com;ksteward@bacbarcodes.com;;;;;
Subject:
rep-PCR data
analysis |
Dr. Renter,
I am responding to your request to M. Healy for advice
on analyzing rep-PCR
fingerprint data. My own background is in
statistics, and I have been
working with this data for two years. This
email outlines my view of the
problems and possible solutions.
You
are correct in recognizing that your question is different from our usual
approach. Our primary interest is in inferring genetic relationships,
while
your question concerns relating genetic relatedness to non-genetic
factors.
>> Yes exactly
The difficulties in
doing this with rep-PCR data are simply those common to
all analyses of data
sets of high dimension.
The high dimension of the genetic data is
essential our concern
Is
modelling correlation coefficients a concern?
We have not worked out the
answers yet, but we
have thought about the issues.
The aim is to explain the variation
observed in rep-PCR fingerprints in terms
of various risk factors.
This calls for an analysis of variance (ANOVA)
approach. ANOVA
accounts for the total variation in a response variable (the
rep-PCR
fingerprint, in this case) by quantifying the degree to which each
covariate
explains the outcome.
>>
This is essentially what we want to do, although we were planning on using a
regression analysis rather than an ANOVA. We could use a linear, logistic,
polynomial, etc. regression depending on what is appropriate (and why we asked
about the 'data structure'). A colleague of ours has suggested a 'matrix
regression' and he implied that the matrix could constitute the dependent
variable. We are not familar with this approach and have not been able to
find references to this method (and are not sure it exists or would work for
this application).
You point
of Normality below is one reason we would avoid a 'regular' ANOVA. In
addition, we would like to describe the magnitude of an association rather than
just test for significance (e.g. calculate Odds Ratios for risk factors that are
significant in the final model). Also keep in mind that due to the
observational study design and sampling scheme we would have a very unbalanced
number of observations for some risk factors.
Several issues keep us
from a simple application ANOVA to rep-PCR fingerprint
data. First,
ANOVA is typically applied to univariate, or low dimensional,
data, while
each of our "data points" is of dimension 850. We can avoid this
problem by using the pair-wise distances as the (univariate) response
variable. This, however, would introduce the complication that the set
of
pair-wise distances violates the important assumption of mutual
independence. Also, reliable interpretation of the simple application
of
ANOVA requires that the response variable be Normally distributed (or, at
least in the neighborhood of Normality).
There are several possible
approaches. First, I refer you to a 1992
paper "Analysis of Molecular
Variance Inferred from Metric Distances among
DNA Haplotypes:
Application to Human Mitochondrial DNA Restriction Data" by
Excoffier,
et al., (Genetics 131: 479-491). This paper describes a method of
assessing the significance of hierarchical population divisions based on
pair-
wise distance data. Significance is determined by a permutation
test rather
than by the traditional F statistic in order to avoid relying on
inappropriate assumptions. This method is implemented in the
widely-used
genetic data analysis program "Arlequin"
(http://lgb.unige.ch/arlequin/).
There may be more recent
implementations. An internet search for "AMOVA"
yields a variety of
resources.
The application of
permutation tests is interesting and may be relevant. However, I'm not
sure how we would include (and deal with dependency in) multiple explantory
variables (see variables below).(ie Model vs testing for
associations)
Excoffier, et
al., are concerned with spatial structure, so all explanatory
variables are
binary, and different variables have a hierarchical, rather
than additive,
structure. Their method would require some modification to
apply to
more general risk factors.
We
would have hierarchical risk factors as well (e.g. some herds have multiple pens
within herd and multiple animals within pens). Most of are explanatory
variables are categorical; except perhaps for temporal and spatial variables.
We are familiar with methods that deal with these types of explanatory
variables in regression analyses.
Although there is an individual-based model underlying the Excoffier
method,
the input data is pair-wise distances. Another approach is to
apply ANOVA to
individual-based statistics.
This requires dimension reduction. (This is
your suggestion of "a measure
(or series of measures) for each isolate
that capture the data".) The
challenge is to find statistics that
reduce the information available in a
fingerprint to just one, or a few,
numbers.
Yes
exactly
One very simple way to reduce the dimension is to represent
each fingerprint
by its distance from a "standard" profile. (This, I
think, is your idea of
a "rooted tree".)
Yes.
And if we used this option, we thought it may be useful to do a
sensitivity analysis of the regression model (quite common in epidemiologic
analyses of observation data). In other words, develop a regression model
for risk factors using a 'standard' or 'reference' strain to define the
dependent variable relationship ('genetic relatedness'). Then change the
strain used for comparisons to another 'standard(s)' which would change
(slightly?) the dependent variable and may or may not impact which risk factors
are significant. This would tell us whether or not the 'standard' chosen
for the genetic comparison effects the model (effects the significance/magnitude
of explanatory variables).
If all
fingerprints are small variations of a single type
this might work well.
If a few types appear in the data, you might choose a
few standard
types and represent each fingerprint by its distances to each
standard.
In this case, the data would no longer be univariate and a
multivariate analysis of variance would be appropriate.
A more common
method of reducing dimension is to compute the principle
components
decomposition of the data set, and to choose the first few
components for
analysis. This method optimally (in some sense) represents
the
variation in the data in the fewest possible orthogonal
components.
It seems like some
method of dimension reduction for the dependent variable may be necessary.
I mentioned to Dr. Healy that we would require their input if this is done
so that the 'outcome' variable still accurately capture the genetic
relationships. Perhaps a principle components approach would be most appropriate
for analysis, but I am not sure.
In any case, once the dimension is reduced to a
manageable number, it will be
important to apply a transformation to the
distances so that their
distribution
is approximately Normal. After this, the application of ANOVA
is
direct.
Finally, let me refer you to a recent paper on a similar topic,
Blackwood, et
al., (2003) "Terminal Restriction Fragment Length Polymorphism
Data Analysis
for Quantitative Comparison of Microbial Communities",
published in Applied
and Environmental Microbiology vol. 69, number 2.
The statistical analysis
is not deep, and it does not address your
question directly, but it might
serve to generate some ideas about possible
analyses.
In the above paragraphs I have addressed the problem of
inferring the
importance of risk factors. In your email you go on to
say that you would
like to "group isolates based on the significant
factors". I'm not quite
sure what you mean by this.
What I meant was: if our statistical
model suggests that factorX is associated with genetic variability, then we
would want to create dendrograms grouped by each level of factorX (for figures
in a publication lets say). In other words, only create dendrograms after
the statistical analysis (i.e.we know which explanatory variables are associated
with genetic variability)
More generally, we would like a better idea of
your final aim for the
analysis. It could be that our existing analysis in
the DiversiLab
software may address your need.
Our final aim for the analysis is to
apply 'standard' epidemiologic methods for statistical analysis of observational
studies to these genetic data. That means determining which 'model' best
describes the variability in the genetic data given the explanatory variables we
have collected. This will tell us which factor(s) best fit the data and
therefore the effect (significance and magnitude) of these factor on genetic
variablity.
The main challenge
as we see it is defining a dependent variable that captures the genetic
relationships and lends itself to this type of statistical methods.
Alex Renwick
P. S. As for
the data itself, we are currently exporting it from the
Bionumerics
software. Is there a particular format (e.g., tab delimited)
that you
prefer?
> ---------- Forwarded Message -----------
> From:
David.Renter@gov.ab.ca
> To: tbittner@bacbarcodes.com,
mhealy@bacbarcodes.com,
> kawillia@epi.umaryland.edu
> Sent: Fri,
12 Dec 2003 14:01:08 -0700
> Subject: Data & Analysis
>
>
Hello all,
> John and I have tried to outline our intentions (FYI - by
risk
> factors, we mean the species, farm, geographic location, time,
etc.
> from where the isolates were recovered.):
In general, we
> would like to see which risk factors explain the
genetic relatedness
> among isolates; rather than compare the genetic
relatedness among
> isolates with different risk factors. In other
words, instead of
> generating hypotheses in advance about genetic
relatedness that we
> test (e.g. isolates within farms are more/less
similar than between
> farms) we would like to use the genetic data tell
us which factors
> are important (e.g. does the 'farm' variable explain
relatedness).
> The difference between the two approaches may seem
subtle, but
> statistically these are quite different.
The statistical
> model would be 'genetic relatedness'
=(i.e. is explained by)
> factor1 + factor2, etc. Then we
would like to group isolates based
> on the significant factors (e.g. if
'farm' is significant, then
> build a dendrogram grouped by farm).
In order to do this we
> must first determine the
appropriate data structure that defines the
> 'genetic relatedness'
variable. We can't model a dendrogram
> statistically, but we could
model the data that defines the overall
> dendrogram. The matrix of
correlations could be one data structure,
> but these numbers are not for
individual isolates, but pairs of
> isolates. The risk factors are
for individual isolates (not pairs).
> Can we define a relevant
outcome(s) individual isolates? Perhaps a
> measure
> (or
series of measures) for each isolate that capture the data that
> defines
the correlation matrix. Perhaps a 'rooted' tree would serve
> this
purpose so that all study isolates are compared to the same
> reference.
There are probably other options as well.
>
> I hope this at least
gives you an idea of our proposed
> approach/needs and can serve as
starting point for our upcoming discussion
> Cheers,
> Dave
>
> David G. Renter, DVM, PhD
> Veterinary Epidemiologist
>
Agri-Food Systems Branch
> Food Safety Division
> Alberta
Agriculture, Food and Rural Development
> Phone: 780/422-0275
> Fax:
780/422-3438
> ------- End of Forwarded Message -------
>