Showing posts with label science. Show all posts
Showing posts with label science. Show all posts

Monday, January 28, 2008

Mr. Idiot shows polarity affects isotropic skew

In a comment, Mike pointed us to observation that the isotropic skew is dependent on the polarity of the column, which we know to be different in the Landis GCMS and IRMS runs. He also pointed us to the reference, which we include a snapshot of below, because there is a lot interesting there in very little space.

Note particularly the observations that overlaps are trouble, peak and baseline separation are important, and that Brenna's software isn't used in practice.



Handbook of Stable Isotope Analytical Techniques, Volume I, parts of pages 168 and 169. Click for bigger.

This is a survey article by one Wolfram Meier-Augenstein, and is citing J. Brenna.

Small world.

Full Post with Comments...

Idiots Return to Brenna and WM-A, Part II

Alasdair, once Ali, continues...

In Part 1, we managed to recreate Brenna's results. In Part 2, we're going to explain exactly how they occurred with reference to both the m45 and m44 signals and what happens when peaks overlap. However, before we do that, we need a little more background on what the software does.

Some of you may be thinking "We understand the logic, but m45 leads m44,not the other way round, you idiot."

Well, that's true, but hold on...

[MORE]


Brenna stated that his software compensated for the misalignment between m44 and m45, so we know that an attempt was made to align them perfectly.

That never happened though, because if it had, he would not have observed the results that he did. There are many factors which would make perfect alignment extremely difficult, including: The m44 and m45 signals are sampled quantities; In practise, they are not perfectly the same width and the same shape; The m45 peak is 100's of times smaller in magnitude than the m44 peak and more susceptible to chemical and electrical noise.

So Brenna's software tried but failed to achieve PERFECT alignment.

Older software may not even try to compensate for the m44-m45 time difference. Newer software may try but fail, resulting in a small finite misalignment either one way or the other. Consider this: To achieve perfect alignment and avoid errors when peaks overlap, you need perfect peaks.

At the end of the day, whether your software compensates or not, it looks like you're going to be left with an error in the alignment of m44 and m45. We don't know whether the OS2 software used for Floyd's Stage 17 compensates or not. We don't know if it behaves like Brenna predicted or whether it behaves like Meier-Augenstein predicted. After we've explained the results, we'll tackle that question.



Before we get into what can go wrong, lets see what is required for accurate CIR measurement. Figure 8 shows a m44 peak, with its corresponding m45 component superimposed on it. The two peaks are perfectly aligned. The integration limits are shown as two vertical lines. Remember, m44 and m45 exist as two separate signals, and we need to show both to explain what’s happening. The software needs to calculate the area under the m45 peak, calculate the area under the m44 peak and divide the m45 area by the m44 area. The result is the CIR for this substance.

Let’s say the m45 signal leads (occurs before) or lags (occurs after) the m44 peak by some small amount. As long as there’s no other peak interfering with our peak, that’s not a problem. We just position our integration limits so that they encompass both peaks and we’re good to go


Figure 8. m44 and m45 perfectly aligned. No overlap




Right, lets introduce another peak to the left of our first peak. Figure 9 shows two peaks of the same CIR overlapping by some arbitrary amount. We use the same method as Brenna to set up our integration limits. Lets examine what’s happening with Pk2 . We can see that our integration limits exclude the left hand tail of both the m44 and m45 peaks. That means that we’ve lost those areas from our calculation. That’s not a problem, because the m44 and m45 are perfectly aligned, so we lose the same relative proportions of each.

We’re also including the right hand tail of Pk1. Again that’s not a problem because in this case, both peaks have the same CIR so the contribution from Pk1 has the same relative proportions of m44 and m45 as Pk2 (it wouldn’t matter if Pk1 was much bigger than Pk2, the previous statement holds true).


Figure 9 m44 and m45 perfectly aligned. Overlapping peaks

So with perfectly aligned m44 and m45, it won’t matter how much those peaks overlap, because they have the same CIR.




Figure 10 shows our m45 leading the m44 by some amount (we haven’t labelled the peaks again as they are clearly identifiable from Fig.9). We’ve had to exaggerate this so that you can see what’s happening. What you’re seeing is m45 leading m44 by an enormously large 1 second. [To put that in context, if Figure 4 had used a 1 second lead instead of a 0.1 second lead, the measured value for that peak (which has a true value of –27) would not have been –28.4, it would have been –41.0 !]

We don’t need to look too hard to see that the m44 situation remains the same. We lose the left-hand tail of Pk2, but gain an equal amount of the right hand tail of Pk1. Total area same as before. It’s all going wrong with our m45 peak though. We’ve lost more of our Pk2 m45 than in Fig.9 and we’re gaining less of Pk1’s m45. That’s upset the balance. When we calculate our areas for Pk2 now, the ratio will make it look like we have less m45 than before. That makes our d45 figure more negative and it makes it look more probable that doping was involved.


Figure 10 m45 leads m44. Overlapping peaks




Figure 11 shows the opposite situation, m45 lagging m44. Again, all’s well on the m44 front, no change again.

So what’s happening with our m45 now?

We’re losing less m45 from Pk2 than in Figure 9 and we’re gaining more of Pk1’s m45. That’s upset the balance again, but this time it’s going to increase our m45 areas. That makes our Pk2 d45 figure less negative and it makes it look less probable that doping was involved.


Figure 11 m45 lags m44. Overlapping peaks




There you have it. The mystery of the overlapping peak and why the observations of Brenna and Meier-Augensten were both valid. It depends on the software system you’re using and how it performs with regard to aligning m44 and m45.

All very interesting, but what about the 64,000 dollar question … What is the LNDD system like?

The reprocessing of the Stage 17 EDFs provides our only real insight into the characteristics of the LNDD equipment. Specifically, we’re going to look at the results obtained from the blank sample. This is useful for us because: We know it shouldn’t test positive; It was processed in an automatic mode by the software; It was also processed in a manual mode by the technicians.

The F3 for the blank showed a degree of overlap between the 5bP peak and the 5aP peak (not to the extent that Floyd’s overlapped, but it was there).

In auto mode, the 5aP for blank came out as delta –3.65 (“delta” refer to the numerical difference between the d45 value for our substance under investigation and the d45 value of a reference substance generated by the athlete which is known not to be effected by the doping in question). The threshold for a doping positive is delta –3. Anything more negative is considered indicative of doping, however the uncertainty that LNDD claim is +/- 0.8 so more negative than delta -3 is suspicious, but delta –3.8 is the absolute threshold. As you can see, the known clean blank escaped testing positive only because it was within the uncertainty claimed by LNDD. That was when processed automatically by the software, without intervention. When LNDD redid it using their manual method, they got a result of delta –1.87. It is worth noting that they knew this was the blank and that it should not test positive.

So what do these results tell us?

Suppose the LNDD system behaved like the one Brenna used in his research and made the measured d45 appear less negative than it really is. That would imply two things. Firstly, the blank sample was a "positive", and really more negative than the measured delta –3.65. Secondly the LNDD manual method made the result worse, not better – they moved it in the wrong direction, away from the true value. (It also provides a stunning example of the influence of manual methods to achieve desired or expected results.)

This argues that we can’t accept that the LNDD equipment had the same characteristics as Brenna’s with respect to alignment of the m45 and m44 signals. If it did, the blank must have come from someone who had been doping with testosterone and the LNDD manual adjustment method is laughably inadequate.

Let’s consider that the LNDD system achieves perfect m45 and m44 alignment. Apart from the fact that it’s probably not achievable, it would mean that the blank, non-doping sample only just escaped a false positive. What does that say about the threshold of delta –3 ?. Also, it would mean that the LNDD manual adjustment method changed a correct result to become incorrect by more than delta –1.6. That’s way beyond their claimed uncertainty figure of +/- 0.8 and would call their competence into question.

We’re only left with one other option, and that is that the LNDD system results in the m45 leading the m44 by some finite amount, resulting in the 5aP appearing more negative than it really was. That would explain why the blank almost tested positive. That would explain why LNDD manually altered the result for the blank to make it look less negative. It would also explain why all of Floyd’s 5aP results appeared far more negative than they actually were.

Far more negative? We’ve established that the error is proportional to degree of overlap. Look at the blank F3 chromatograms. The degree of overlap between the 5bP and the 5aP is negligible, but that was enough to make the blank almost test positive. Now look at Floyd’s F3s.

Noticeably more overlap.

For those calling us idiots, let's remember that we are not talking about the actual lag between the peaks, but lags produced when the software attempts to account for the actual lag and fails to to a perfect job. Such software more or less works adequately when there are no overlaps in peaks. As we've shown, when there are minor errors in this compensation, overlapped peaks can be systematically measured with inaccurate results.

We've shown here that it is possible to reconcile Brenna's study and WM-A's theory in a way that leaves it less likely that LNDD got correct results.

Which seems like a perfect place to stop.

Full Post with Comments...

Sunday, January 27, 2008

Idiots return to Brenna and WM-A, Part I

Alasdair has returned from his Hermit's cave having shed his 'Ali' persona, and now revisits some key points in the battle of the experts.

Forward to Part II.

Let's pay a return visit to Brenna’s 1994 paper (Curve Fitting for Restoration of Accuracy for Overlapping Peaks …), which we first looked at in November.

Why are we doing this?

Well this paper formed the basis for a clear difference of opinion between Brenna and Meier-Augenstein during the hearing. All of Floyd’s IRMS F3 chromatograms showed a degree of overlap between the 5bP and 5aP peaks. If this resulted in an unknown error and that error biased the result in a more negative direction, it would/should have put those results in question (more negative equates to higher probability that testosterone was not generated by the athlete’s body).

Meier-Augenstein argued that the error was of unknown magnitude and in a more negative direction. Brenna argued that the error was very small and in a less negative error (i.e. the error actually made it look less probable that the athlete doped).

Meier-Augenstein’s opinion no doubt comes from his years of work and research in this discipline. We can trace Brenna’s opinion back to a paper he published which, amongst other things, detailed his observations when peaks of almost identical carbon composition overlap. His observations were that a systematic error occurs when peaks overlap and that error (not small, by the way) made the CIR for the substance being investigated look less negative than it really was.

What he didn’t investigate in that paper was why that happened. Understanding results is usually a prerequisite to applying them to different situations. Brenna chose not to take that precaution in his testimony.

[MORE]


Enter the idiots, tattered old spreadsheet in hand. We decided to finish this job off by solving the mystery of the overlapping peak and attempting to prise Meier-Augenstein’s fingers from Brenna’s neck.

Let’s start with a review of this test.

The GC/C/IRMS process separates the C12 and C13 isotopes and produces two individual signals, the m44 and the m45 peaks (corresponding to the C12 and C13 isotopes, respectively). The relative size of these two signals is proportional to the carbon isotope ratio (CIR) of the substance under investigation. The CIR is actually calculated as the area of the m45 peak divided by the area of the m44 peak, using integration. In an ideal world, the m45 peak can be thought of a scaled down version of the m44 peak (100’s of times smaller). In practice, these two peaks do not align on the time axis perfectly, the error is very small (typically 150 milliseconds for a peak that may be approximately 35 seconds wide at the base).

Brenna mentions in his paper that his software corrects the misalignment of m44 and m45. He also states that he uses the “perpendicular drop” method to separate peaks (a vertical line positioned at the minima of the intersection of the two peaks and drawn down to the background level). Can our spreadsheet cope with these demands?. Darned right it can !

The goal of this first part is to recreate Brenna’s observations, so let’s dust off our spreadsheet and fire her up. We’re going to define two peaks, each with a d45 value of –27. We’re going to set the time difference between m45 and m44 to exactly zero and then make sure we’re measuring them accurately. Figure 1 shows the output. Two peaks, no overlap, both reporting a correct d45 of -27. [note: d45 defines the CIR of a substance relative to an international standard in parts per 1000, so –27 reads as a CIR which is –27 parts per thousand less than the standard].


Figure 1. Exact m45 and m44 alignment. No overlap between peaks.

It’s worth noting here that we only show the m44 trace, as appears to be common practice when presenting these results, but remember that there’s another trace (the m45) not shown.

Right, let’s introduce some overlap. Figure 2 shows the results we get. Note that we’re still measuring the true value. No error there.


Figure 2. Exact m45 and m44 alignment. Moderate overlap between peaks.


Let’s increase the overlap. Figure 3 shows a big overlap. Unfortunately, there’s still no error. What’s happening here?



Figure 3. Exact m45 and m44 alignment. Significant overlap between peaks.

OK, in our ideal world with two equal peaks, both having perfect peak responses and perfectly aligned m45 and m44 responses, it would appear that overlap has no effect on our results.

The question now, is what factors were at work to produce the results that Brenna observed?

Let’s go back to the drawing board. We know that m45 tends to lead m44 so we’ll scrap our assumption of perfect peak alignment and allow the m45 response to lead the m44 response by some nominally small value, say 100 ms.

Figure 4 shows our two peaks again. Same situation as Figure 2, but with the m45 response leading the m44 response. Now we’re seeing something change. The measured value is wrong and it is more negative (remember, the true value of our peaks are both -27)


Figure 4. m45 leads m44 by 100ms. Moderate overlap between peaks.

Figure 5 shows an increased overlap. Same situation as Figure 3, but with the m45 response leading the m44 response. The overlap has increased and so has our error. This is starting to resemble Brenna’s research, apart from one major difference – our error is going in the opposite direction to his observations.

Having the m45 response lead the m44 response gives the results that Meier-Augenstein predicted – an error which makes the results appear more negative. That was also his stated reason for the response he predicted. He must have observed systems where the m45 signal lead the m44 signal by some finite amount.


Figure 5. m45 leads m44 by 100ms. Significant overlap between peaks.

Right, let’s review what we’ve observed so far. In our ideal world, with no other sources of error other than the alignment of the m45 and m44 responses, perfect alignment of m45 and m44 results in no error when peaks of identical CIR overlap. We can infer that Brenna could not have had perfect alignment of the m45 and m44 signals. Having the m45 lead the m44 results in a systematic error, proportional to degree of overlap, which biases the result in a more negative direction.

Let’s see what happens if we have m45 response lagging the m44 response by some small amount, again 100ms. Figure 6 shows the same situation as Figure 4, but with the m45 response lagging the m44 response. Right, now we’re getting somewhere. An error which makes the result less negative. So far so good.



Figure 6. m45 lags m44 by 100ms. Moderate overlap between peaks.

Figure 7 shows an increased overlap. Same situation as Figure 5, but with the m45 response lagging the m44 response. The overlap has increased and so has our error. This looks just like the results Brenna observed. A systematic error which is proportional to degree of overlap and results in a less negative bias.


Figure 7. m45 lags m44 by 100ms. Significant overlap between peaks.

OK, so we can provide an explanation which satisfies both Brenna’s observations and Meier-Augenstein’s predictions. It’s all to do with the relative positioning of the m45 and m44 responses. Perfect alignment and there’s no error. A small misalignment (in our case, approximately 0.3% of the peak width) and a systematic error is observed proportional to degree of overlap and biased in either the more negative or less negative direction (depending on whether m45 leads or lags m44).

This is probably a good point to establish some ground rules. Perfect alignment is unachievable. This is not an analogue system with continuous signals. It is a digital system which samples the m44 and m45 responses at some fixed frequency. Regardless of what algorithm is used by the software to try and align the two signals there will always be a finite error in that process. There are many other factors which make perfect alignment of m45 and m44 an impossibility in practice, e.g. slightly different peak widths, slightly different peak shapes, etc. All these factors conspire to make the job of the software a demanding one.

The next part will demonstrate why we observe these results and what we can infer from the LNDD results.

Forward to Part II.

Full Post with Comments...

Tuesday, January 01, 2008

Mr. Idiot figures something out

In a comment to "Getting around to Specificity", Mr. Idiot reaches some important understandings, and we post it for broader visibility. It's been a real slow couple of weeks, and what comments there have been are in topics lost down the page.


My basic understanding is now that LNDD's method of detecting the presence of exogenous testosterone is not "fit for purpose." This failure of the assay to be fit for purpose has come about because LNDD failed to take account of a key difference between the pre-IRMS testosterone test, and the GC/C/IRMS testosterone test.

Before IRMS (about 1997 at UCLA - I don't know exactly when LNDD got it, but certainly before 2003), testosterone was clearly a threshold substance. The T/E ratio had to cross the quantitative threshold of 6/1 (now 4/1). The key measurement was the AMOUNT of testosterone compared to the amount of epitestosterone.

In that context, TD2003IDCR makes perfect sense. In order to use GC/MS to detect exogenous testosterone you can identify and measure the amount of testosterone and epitestosterone with a GC/MS in Selected Ion Monitoring (SIM) mode. By monitoring three ions of each substance you are able to both clearly identify and quantify both T and epi-T. First you separate the substances with the GC, and then monitor the three key most abundant ions (i.e. the "diagnostic ions") of the proper GC peak with the MS.

The key thing to note here is that even if your GC separation is not perfect, or even if there is an indistinguishable perfectly co-eluting peak, leading to other substances in your GC peak, it does not invalidate your measurement. If the ratio of the three diagnostic ions is right (those are known ratios) then you know you have the right substance with no interference AT THOSE IONS. There maybe other things in that peak, but, when you are not using IRMS, it doesn't matter because you have identified and quantified your testosterone properly with the three diagnostic ions. Same for epi-T.

[MORE]


(This is not really relevant to my present topic, but as it turns out LNDD only analyzed ONE diagnostic ion in their T/E test, so the arbs ruled their process violated TD2003IDCR and that the evidence of that test was without value. If that was the only test, the case would have been dismissed.)

Now, when TD2003IRCR was approved in 2003 IRMS was still relatively new, and not all WADA labs had an IRMS machine. When TD2004IAAS became effective in August of 2004, IMRS was still officially an "add on" - that is, if the T/E confirmation test, describe roughly above, was not conclusive, then an IRMS test was recommended. Thus TD2003IDCR was written so that administration of exogenous testosterone could still be confirmed without IRMS - and, in fact, as best I can tell, that is still the situation today. That said, even WADA understands that the IRMS test is much more conclusive when done properly.

So, historically speaking, you have the situation in which GC/MS works "fine" (at least by WADA standards), and then, at different times in different labs, the IRMS test is added as an additional piece of evidence.

Now, what LNDD apparently did when they added the IRMS test to their arsenal, was to continue to do the GCMS test the same way they had before - with SIM - the way it had been perfectly effective before, and, importantly, a way it is perfectly in line with TD2003IDCR.

HOWEVER, when the IRMS test is added to the GCMS test, the GCMS test must be done a little differently. Of course, you are analyzing metabolites of testosterone, not T itself, but that is not important here. The key thing is that SIM no longer does the job you need done. For GC/C/IRMS you not only need to know that there is substance x in a given GC peak, but you also need to know that there is nothing else in addition to substance x in the peak. SIM mode does not tell you this. Only doing a full ion scan of a given peak will tell you this. That is, you have to monitor not just three diagnostic ions, but all ions present in the peak because all the ions are going to the IRMS. Remember for the T/E test, it didn't matter if other stuff was present. The SIM/three diagnostic ions told you everything you needed to know. But doing the same thing as part of the IRMS test is not good enough. It does not give you the information you need.

So, if the GCMS part of GC/C/IRMS is done in SIM mode then the whole assay is not fit for purpose, even though, for "historical" reasons, it is probably technically allowed according to TD2003IDCR.

The wise reader will ask why WADA has not yet recognized and fixed this problem. Well, I can only guess that it is because, prior to the Landis case, no lab had run the GC/C/IRMS assay in exactly this way. You see, while TD2003IDCR says that SIM mode is acceptable, it also says the "preferred method" is to use full scan mode. And if you run the assay with this full scan mode then you have the conclusive evidence you need that you have measured the right stuff in the IRMS. I imagine that other labs have either not done IRMS or have done IRMS with the GCMS part done in full scan mode. Dr. Goldberger testified at the hearing that he had seen lab documentation packages from the UCLA lab and they included the full scan data for the T/E test - and if they did it for that, they would surely do it for the GCMS part of the GC/C/IRMS test also.

So, why didn't Landis' legal team press this particular issue at the hearing - that the way LNDD does the GCMS part of the IRMS test makes the assay not fit for purpose? Well, one answer is that, as Larry said above, that is a very difficult legal argument to make given the nature of the controlling documents. Especially given that the way LNDD does it appears to be okay according to TD2003IDCR. The truth is that TD2003IDCR has inadequately accounted for the nature of the IRMS test - leaving the "loop hole" of running SIM mode open, even when doing IRMS.

Another answer as to why Landis' legal team didn't press this at the hearing is to say that "They did, sort of." This is exactly what they were trying to get at with all the arguments about good chromatography. If they could not use (the inadequate) TD2003IDCR to prove an ISL violation, they had to get at the issue another way. If they could show that there was a some degree of likelihood of other material in the peaks of interest, then that should have raised enough questions to prove an ISL violation of the "matrix interference" bullet of ISL 5.4.4.2.1. But there are two problems with that approach - ISL 5.4.4.2.1 has weak language ("should" rather than "shall"), and also that there are no cut and dried criteria for what constitutes good or bad chromatography. The arbs, who did NOT understand the critical flaw in applying TD2003IDCR to IRMS, allowed the use of SIM to stand, with mediocre chromatography, because "should" does not mean "shall."

And this is where the "threshold substance" vs. "non-threshold substance" argument comes in. The way the IRMS test works, clearly testosterone SHOULD be a threshold substance, and the stronger language included in ISL 5.4.4.2.2 should apply. But because WADA has not yet admitted the fatal flaw in applying TD2003IDCR to the IRMS test, both Landis' legal team and the arbs assume that testosterone is a non-threshold substance and the weaker language of 5.4.4.2.1 applies.

There is more to talk about with regard to what exactly could be in those peaks of interest other than what should be there, but that's obviously enough for now.

Bottom line is that LNDDs method of using the SIM mode for the GCMS part of the GC/C/IRMS test makes the assay not "fit for purpose."

Oh, I just have to add one more thing - about ISO certification. This apparent technical adherence to TD2003IDCR is why LNDD could get ISO re-certification for its IRMS test just six months before Landis' tests. ISO certification does not assure that the test really does what it is supposed to do. It only assures that the test is in line with the controlling documents. If those documents are flawed, that's not ISO's fault.

SYI

We've noted before that the WADA system appears to have no body responsible for "fitness for purpose" review of test protocols. If a dunking stool were used for determination, and it's execution were to the letter of the SOP, ISO would be OK with it.


Full Post with Comments...

Saturday, December 29, 2007

Apples and Oranges

From a dubiously respected source, we find proof of the validity of the Visual Gestalt ("eyeballing") method of matching chromatograms.


Figure 1: Infrared transmissibility curves

By the Brenna criteria of general arrangement and relative heights of peaks, there appears to be complete isomorphism between the targets, showing that one can indeed compare apples with oranges.

Full Post with Comments...

Wednesday, December 19, 2007

TIC...TIC...TIC...TIC

Is there something between the 5bA and 5aA? Do the squashed TIC graphs tell us anything? Click images for full-size.



Fig 1: A sample F3 Blank TIC, reported values 5bA -27.54, 5aA -28.40

Fig 2: A sample F3 TIC, reported -28.82, -32.12

Fig 3: B sample F3 blank TIC, reported values -27.54 and -28.31

Fig 4: B sample F3 TIC, reported values -28.79, -31.88



Fig 5 summary: A blank, A sample, B blank, B sample, identically scaled


By appearances, the A blank looks clean, the B blank has a minor signal, and both of the Landis samples have a noticable and distinct peak between the 5bA and the 5aA. This does not appear in the graphs of the individual peaks as one of the three ions for the expected 5bA and 5aA, suggesting is is something else that has not been accounted for.

How big is the intervening peak?



Fig 6: scaled 0 to 7 million in 350 pixels. Minor peak is 11 pixels high, or 3.14% of the 5bA, or 6.28% of the 5aA.


(Initial observation by SYI in comments elsewhere).

UPDATE en route: we're told the graphs are better on other pages; we'll update when possible.

Full Post with Comments...

Monday, December 17, 2007

Three kinds of display

M provided a link to a page at the American Society for Mass Spectrometry that included the following very useful illustration identifying the type of displays that can be provided.

We see identified three kinds of traces.

  1. Total Ion Count (TIC) reflecting the counts of ions of all masses for the entire run.
  2. Selected Ion being monitored (SIM) the count of particular ions at selected masses for the whole run
  3. Mass spectra for individual peaks, in this case we think a TIC for each peak.
Not shown are SIMs for individual peaks, or traces of individual ions over time, which we have also seen.

As a sometime programmer, I am likely to dump the collected data into a matrix (or a spreadsheet), with one dimension reflecting sample times, and the other reflecting the ion in question. For instance, using a spreadsheet, one might have 500 columns reflecting m/z 50 through 550, and however many rows are needed for the sample interval.

One can imagine the possibility of higher achievable sample rates with fewer ions being collected, a full scan offering perhaps less resolution than SIM collection.

Given the complete data, the selection of a particular display is an issue of presentation, not of the data collected.

Storage

As rough estimates, assume a 32 bit value (4 bytes) for each sample of 500 ions, giving about 2k bytes of memory per sample. We have about 2600 seconds in a run, and we have isotropic offset of about 150ms. That suggests to me a sampling rate of at least 150/3ms, or 50ms which is 20 samples per second, or 52,000 samples on a 2600 second run. At 2k per sample, that is about 100MB per run.

This kind of data is highly compressible by at least a factor of 10, so I'd expect the complete data for a run to be less than 10Mb. About 70 of these would fit on a CD-ROM. (I'd really expect compression of 100 or so, since most ions will all be zero.)

Full Post with Comments...

Friday, December 07, 2007

DailBob on different columns

In a belated comment on Different Columns detailed, DailBob does some 'splainin that seems worth fuller exposure:


Before I start, a reminder that we don’t have an IRMS in our lab, so I’m “projecting” some of what I’ll call “GC 101” identification principles onto the IRMS. I think I’m correct in doing this, because this is essentially what the Landis team did via their retention time arguments. I also vetted my thinking with our Analytical Chem Manager, and he’s in the same place I am.

[MORE]



I first want to provide some history and an example to give you some background on what’s driving my perspective. I’m old enough that I remember getting our first GC-MS in the late ‘80s. The addition of the MS was a game changer, because it added another “layer” of identification, the value of which can’t be understated. Prior to that, you could only identify by elution time, and we were sometimes wrong. Now, we’re rarely wrong. This changed the analytical chemistry world to where I think most Chemists would tell you that they require both for identification, i.e., the MS upped the bar. I’m telling you this, because not having the full mass spectra to aid in identifying something on the IRMS places all the emphasis on elution time.

A few weeks ago, we got a sample in where one of our plants thought they accidentally put glycerin in a batch instead of what was supposed to go in. To determine if this happened, we ran a glycerin standard which eluted at 2.854 minutes. We then injected the sample, where a peak eluted at 2.862 minutes (0.008 minutes or about half a second difference). The mass spectra matched the glycerin standard, so these two pieces of information, taken together, identified the material as glycerin. I’ve provided this example to demonstrate how these machines identify compounds. This is the only way they identify compounds, and the identification is pretty “black or white” – either it’s identified, or it’s not (an exception is isomers, but that’s NA here).

I think the main issue as you move a sample from the GC-MS to the IRMS is maintaining the identity of the compound for which you want to determine the CIR. I say this for a couple of reasons. First, even if the GC-MS and IRMS use the same column, the fact that the columns are two separate physical units can cause differences in elution time, or even peak swapping, if one of the columns is contaminated. Second, as I mentioned, because the IRMS cannot identify the compounds with full mass spectra, the only way of identifying it is by elution time. Consequently, I think it is absolutely paramount that it be tied directly to a known standard, because without it, you have zero “layers” of identity.

I have expressed concern about the different temperature ramps between the two machines, which you have discounted (note: Larry has uncovered information that makes it look like the difference in temperature ramps is intentional, but this doesn’t negate my point). If you look at the Landis F3 from the GC-MS, there is a peak that elutes at about 13.2 minutes. This peak seems to have disappeared in the F3 IRMS. This peak either a) contained no carbon, b) moved due to the change in temperature ramps, or c) indicates a contaminated column. I can make this argument for any peak that is in the GC-MS output that doesn’t appear in the IRMS. If I can't account for a peak, and the peak I'm interested in is not directly identified by being tied to a standard, its identity becomes suspect.

You asked me whether I think the 5aA or P could be misidentified? If you can play a mental game for a moment and adopt the perspective that the only way things are identified are by elution time (minimally) and mass spectra (ideally), I think you can see that, technically speaking (i.e., by standard separation science), they aren't identified at all. What you have done in your example is arrived at the identification by logical deduction. While your logic is excellent, logic is not science. Because of the gravity of this case on a person’s life, I have to have the identification be done by standard separation science, particularly because it's capable of doing so in an incontrovertible way, if it's done carefully and correctly.

You asked me once whether I thought the science told us whether Floyd doped, and I replied that I really couldn't tell with any surety. This is why. While I can't prove he didn't dope, this science isn't good enough, in my view, to prove that he did. When I combine this with a) generally crummy chromatography, b) poor lab practices (e.g. white-out), c) the fact that the blank urine was “positive” (for me, indicating some possible machine bias) and d) the fact that I was influenced by Amory’s testimony, I still have enough doubt that the excellence of your logic doesn't push me over the hump.

Full Post with Comments...

The 5bA Anchor argument

One of the arguments lying around is that the 5bA is sufficient to anchor the identification in the Landis F3, based on the following reasoning. The mix-cal acetate has an internal standard, 11k-etio, 5bA, and andro, some of which appear in all fractions. Using the pattern matching method, we claim we can see in the Blank F3 which the IS and the 5bA are matching from the mix-cal acetate. We know the blank also has 5aA, so the peak that is in the blank that matches the peak in the Landis F3 must be 5aA, and similarly for the 5bP. Even though we don't have 5aA or 5bP in the calibration standard, we can extrapolate by transitivity through the blank.

This is shown below:


Figure 1: Formerly Fig 5 of Retention Times II.


The yellow bars are the things we claim match up through all the samples, the mix-cal, the blank, and the Landis F3.

[MORE]

We should note that there are still some open issues about exactly which peak is claimed to be the IS in the Landis F3, because the area is quite cluttered. This is what led WMA to wonder if LNDD decided which one matched by looking at the CIR value of the peak, even though the measured value wasn't entirely within the spec for the CIR of the IS. This adds some lingering uncertainty, because the SOP calls for adjusting things so the IS comes out in the vicinity of 870 seconds, which requires you to know which peak is the IS. If you're off, then later peaks might also be off.

Having claimed to match the IS, and duplicated other conditions, then the aligning peaks in the blank and Landis F3 are taken to be the 5bA.

Figure 2: Claimed matches on the IS and 5bA on the mix-cal, blank and Landis F3.

Now things get transitive. We don't have a standard for the 5aA or 5bP, but we believe they are in the blank. So if we know what they are in the blank, we can similarly match. Let's look at the 5aA first.

Figure 3: The peak claimed to be 5aA in the blank is mapped onto the Landis F3


The identification is being made in the blank based on the belief that the 5aA follows the 5bA -- see Shackleton, who used the same column, for example. But what about that peak that is slightly after the one we're looking at?
Figure 4: There's a little peak just past the one claimed to be 5aA in both of the samples.

We have a small peak in both samples just beyond the one claimed to be 5aA. Could it be the real 5aA? If so, how would we know? We don't have certainty from a calibration mix of any kind. What we do have is a general idea from Shackleton that the 5aA follows closely, and that it often appears the 5bA is tailing into the start of the 5aA. The argument would be that since we have that kind of tailing from the claimed 5bA into the claimed 5aA, that's what they must be.

Leaving open the question: where does TD2003IDCR talk about proximity and tail shape as identification criteria?

It is also very interesting that LNDD took the CIR of the peak following in the blank, but did not in the Landis. Why would they take one measurement, and not the other, for a peak in the same location in both samples? We see in the blank that peak is less negative than the claimed 5aA, and closer to the value of the claimed 5bA. We know it is trivial to take such a measurement, so why wasn't it made for the Landis?

The same argument arises in the 5bP:
Figure 5: Claimed 5bP in the blank, then mapped down to the Landis F3.


There is much less of a peak following the claimed 5bP in the Landis F3:

Figure 6: Peak following claimed 5bP in the blank is minute in the Landis F3


So, there we are. Absent the actual analytes in the mix-cal, we have to extrapolate identities from the blank onto the athlete sample. This isn't a technique specified in TD2003IDCR. The truth of that claim is built on claims that the IS and 5bA were correctly identified in the blank, which leaves us with the previous discussion of the validity of retention times and general pattern matching.

Or, as DailBob notes, the 5bA anchor theory is trying to reach identification by logical argument, not science.

Comment away...


Full Post with Comments...

Sunday, December 02, 2007

What is documented where

Following up the revelation that LNDD used different columns on the S17 when their standard operating procedure (SOP) says they should be the same, we went looking for documentation, and got some surprising results. Here is a summary of our examination.

Date
UCI
Sample
Exhibit
SOP
MS
SOP
IRMS
Actual
MS
Actual
IRMS
7/3
993865
Ex 92
LNDD
1427
LNDD
1453
missing
missing
7/11
994203
?
?
?
?
?
7/13
994277
Ex 88
LNDD
1045
LNDD
1071
missing
missing
7/14
994276
Ex 90
LNDD
1236
LNDD
1262
missing
missing
7/18
994075
Ex 86
LNDD
853
LNDD
879
missing
missing
7/20
995474A
USADA
LDP

missing
USADA
153

USADA
124

missing
7/20
995474B
USADA
LDP

missing
USADA
329

USADA
303

missing
7/22
994080
Ex 87
LNDD
951
LNDD
977
missing
missing
7/23
994171
Ex 84
LNDD
664
LNDD
690
missing
missing
Aguilera
NL1
Ex 89
LNDD
1142
LNDD
1168
missing
missing
Aguilera
NL2
Ex 85
LNDD
758
LNDD
784
missing
missing
Aguilera
NL3
Ex 93
LNDD
1521
LNDD
1547
missing
missing
Table 1: Where column types are identified

Looking at this altogether, here are a few points of note:
  1. All the tests specify MAN-52 and MAN-41 as the SOPs and they call for the same column, a DB-17.
  2. Nowhere is there an indication of what column was actually used for the IRMS, only the SOP pages. We don't know what column was used for any IRMS test from the documentation. With the data now available, it is no longer safe to make assumptions.
  3. The indication of column used for the MS is only present in the samples run on IsoPrime1, there are none for IsoPrime2. We don't know what column was used for the MS on the alternate B samples. We similarly have no reason to trust the SOP and execution agree.
  4. The SOP pages for MAN-52 should have come at the beginning of the IRMS section in the LDP, but we get the "actual" page only.
  5. The "actual" pages for column used should have been present in the alternate B's next to the SOPs, but were not there.
The above suggest that the LNDD does not have a uniform method for constructing LDP's, and just throws in what they think is appropriate for the sample in question. It also appears that what they do provide is inadequate to determine if they actually did what the SOP says they should have done. This makes it impossible for a qualified expert to determine if what they did was correct. We can suppose this comes from one or more of three causes: indifference, incompetence, or intention to be opaque.

If anyone paws through the exhibits and finds things we've missed, please let us know.

Full Post with Comments...

Saturday, December 01, 2007

Different Columns -- Detailed

In Different Columns yesterday, we passed on the breaking news that LNDD did not use the same GC column type on the Mass Spec that they used on the IRMS on every test done in the Landis case. They used an Agilent 19091s-433 on the MS, and a DB-17 on the IRMS. We'll call the one used for the MS the '433 from here on out.

Comments since have wondered whether this

  1. matters at all to results
  2. matters at all to the legal case
  3. late notice is a failure of the Landis defense team;
  4. late observance is indication even Landis thinks it doesn't matter.
As clarification, our conduit, Mr. Idiot, says that this information comes from Arnie Baker, and that Arnie pointed it out to him and it wasn't an independent observation. And Baker flags it as a "case dispositive" item. We think that answers points 3 and 4; it was a late catch by Team Landis, and they think it is important.

(We'd assume it is raised in the filed appeal brief, which we'd like to get released. USADA has it by now, don'tcha think?)

[MORE]


As to point 2, legal implications, let's look at the page that identifies the DB-17 as the SOP column for the MS, LNDD 664 from Exhibit 84.

LNDD 644: DB-17 in the SOP for the GCMS in MAN-52
(click for bigger)


The original LDP (the USADA pages) do not ever contain the method description page for MAN-52, nor was it produced during discovery, though it was requested. It only appeared in the Exhibits for the B sample tests. Was this an intentional omission in the LDP, and an accidental inclusion in the exhibits? We'll probably never know.

The legal implication seems direct. Their own SOP produced as documentation of runs says they should use the DB-17, but the run results show them using a different column. This would appear to be an ISO and ISL violation.

Baker says in mail to Mr. Idiot that USADA 104 also documents MAN-52 as the SOP for the MS to be used in the CIR test:


USADA 104: MAN-52 is the SOP for the MS in the C12/C13 testosterone test.


Baker alludes to COFRAC accreditation documents for the EC31 test identified above as specifying the DB-17, but we do not have a copy of that document to show.

Now, on to question 1, does this make any difference to the results?

The claim is that columns of different types can change the order of elution of compounds. Baker offers in the mail to Mr. Idiot the following support for that claim.

The Agilent website documents the columns and gives reference chromatographs for each. We're shown acenaphthylene eluting before acenapthene with the '433, and in the other order for the DB-17.

'433 Example from Agilent

DB-17 example from Agilent; reverse order of the a'ene's.

Baker notes for nitpickers that acenaphthalene and acenaphthylene are synonyms, and summarizes for us idiots:



HP-5ms '433 DB-17ms
(%-Phenyl)-methylpolysiloxane 5% 50% virtual
Polarity Non-polar Mid-polar
Acenaphthalene vs. acenapthene Elutes before Elutes after


OK, that's one example, but does that prove anything for the testosterone test? We don't know we have any acenaphthalene or acenapthene, so maybe it doesn't matter.

Baker doesn't have indisputable proof of switches with these columns of compounds known to be present in the samples. He does offer numerous examples of switches using similar columns, in particular a set from a drug screening test. He says the DB-5 is about the same as the '433, and pulls this from Agilent



The ordering on the DB-17 vs the DB-5 is:

DB-17
1-2-3-4-5-6-7-8-9-10-11-12-13-14-15-16-17-18-19-20-21-22-23-24-25-26-27-28-29-30-31-32-33-34
DB-5
1-2-3-13-5-6-7-20-10-11-8-14-9-16-17-4-12-18-19-21-25-23-22-24-15-27-26-29-28-30-31-32-33-34

Is this ordering relevant?
There are some huge differences in the ordering above: DB-17 peak 4 is DB-5 peak 16, and DB-5 peak 3 is DB-17 peak 13.

Baker finds a case from Skogsberg et. al with steroid analysis:



Skogsberg, U. et al. Investigation of the retention behaviour of steroids with calixarene-based stationary phases by modern NMR spectroscopy. Journal of Separation Science, vol. 26, p. 1119-1124. (2003).

In this example, steroid analytes 2 and 3 get switched. Ayalyte 1 is norethisterone, 2 is norethisterone acetate, 3 is chlormadinone acetate and 4 is testosterone acetate.

Let's pretend, for argument, that the column switch did not affect the elution order of the 4 primary analytes in the Landis F3, the5aAC, 5bA, 5aA, and 5bP. We don't know that, but let's pretend.

Given the switches seen in the examples above, how can we be sure that any of the other compounds haven't been moved, perhaps a lot, between the columns? Remember the big changes in the DB-5 vs DB-17 example above.

How can we be sure that any of the peaks in the IRMS are properly identified, and that they do not contain co-elutes that have arrived solely from the column shift? It is not clear that even full-scan mass-spec of the MS portion of this test will tell us anything about the peak contents in the IRMS. Given this column change, it would appear necessary to get full-scan MS from the column before combustion in the IRMS, and this wasn't done.

If the change in column is an ISL violation that causes a burden flip, it looks difficult to prove that the reported measurements were taken correctly, unaffected by the ISL violation.

That is why Baker considers this a "case dispositive" error.

Now, he thought the retention time argument was good too, and we saw how that was dodged. We'll have to see if there are holes in this argument as well.

Start shooting.

Full Post with Comments...

Friday, November 30, 2007

Different Columns!? Hiding in plain sight

In comments, Mr. Idiot has noticed something that was in plain sight all along, if anyone had noticed.

The Majority wrote, in the key seven paragraphs, at paragraph 188:

The GC column is, of course, the same in both instruments.

This is a "misstatement of fact" by the Majority, on which some significant conclusions are based.

We see that the columns were NOT the same, based on parts of USADA 124, 153 for the A sample, and USADA 303 and 329 for the B sample:

[MORE]


USADA 124: Using an Agilent 19091s-433 for the GC/MS.

USADA 153: Using a DB-17 (as did Shackleton) for the GC/C/IRMS

USADA 303: Using an Agilent 19091s-433 for the GC/MS.

USADA 329: Using a DB-17 for the GC/C/IRMS


These are completely different columns, with different "polarities". Polarities are known to causes significant location changes in peaks, including changing the order of peaks. It isn't known as fact whether these columns change the order of the 5bA and 5aA peaks in the setup used by LNDD.

For the record, Agilent used to be part of HP, and now also owns the rights to the DB-17. So both are made by the same company, but are of completely different specification.

Mr. Idiot writes:
[T]his all comes from Arnie Baker through email. I do have explicit permission to make it public. The down and dirty version of it is this - most of it is right there in the original LDP:

USADA 0124 says that the column used in the GCMS was an "Agilent 19091s-433." If you dig in the Agilent website, it is clear that an Agilent 19091s-433 is an "HP-5ms" column.

USADA 0153 says that the column used in the GC/C/IRMS was a "DB-17ms." Again at the Agilent website you can find the DB-17ms and learn that it is a different column with different characteristics than the HP-5ms.

Just as one example, the HP-5ms is "non-polar," and the DB-17ms is "mid-polar."

Baker has shown that the LNDD SOP required that the DB-17ms be used for both GC/MS and GC/C/IRMS, but for some (undocumented) reason they used the HP-5ms for the GC/MS. The evidence of that is a little more complicated, so I'll wait on explaining it.

The information Baker has given me also shows specific examples of how substances elute at different times and in different orders on these two columns, and even has an example of a steroid doing so (although not specifically 5aA or 5bA). That's part of the stuff that is graphics heavy and TbV will have to do something with it.

Obviously one of the first things you learn about the connection between GC/MS and GC/C/IRMS is that the chromatographic conditions have to be the same. Now we realize that not only were the temperature conditions different, the columns were different too.


[BACK FROM FULLPOST]

UPDATE: more in Different Columns - Details.

Full Post with Comments...

Tuesday, November 20, 2007

An Idiot Looks at [Brenna 94]

Dr. Brenna was an author on a 1994 paper that has been cited variously in the case, both for and against Landis, by the usual suspects. Contributor Ali has gotten a copy, and files this evaluation...

By Ali


Curve Fitting for Restoration of Accuracy for Overlapping Peaks in Gas Chromatography/Combustion Isotope Ratio Mass Spectrometry

by Keith J. Goodman and J. Thomas Brenna

(Hereafter [Brenna 94])


The purpose of this review is to summarize their findings and highlight any aspects that may have relevance to the matter at hand – the Landis case.

The background of the paper is that overlapping peaks in IRMS analysis result in inaccurate calculation of o/oo values, which we've looked at before. It says,

The conventional algorithm resulted in systematic bias related to degree of overlap

And it attempts to offer a new algorithm involving curve fitting to get better results. It also offers some experimental results on the affects of overlaps, which are of interest to us.

[MORE]


The conventional algorithm is separating the peaks with a vertical line at the centre of the valley between their overlap and taking that line down to what is assumed to be the background level. This is then used as an integration limit.

It's what our spreadsheet does, and appears to have been employed on a number of Landis’s IRMS F3 chromatograms to separate the 5B and 5a peaks from either themselves, or more frequently from some unidentified small peak that appears to be between and overlapping both the 5B and 5a peaks.

The paper gives a brief description is presented on the GC/C/IRMS. An integration time of 0.25 s was specified, which we believe to be analogous to the sampling rate. Peak start and stop are detected using only the 44 plot (due to superior signal to noise ratio). This differs from the LNDD process described by Brenna, where the 45/44 ratio plot is used to determine the start and stop times. Peak maxima are detected on each plot (44, 45 and 46) to identify the time shifts between the three detectors and the previously determined integration interval is applied to all three plots. Background is identified by a straight line fitted between peak start and stop points (which may be at the same level for constant background or at different levels for sloping background).

Due to the extra plumbing involved in the GC/C/IRMS process, both chromatographic efficiency and peak shape are detrimentally affected. In other words, you tend to get more peak overlaps and the peaks are generally not Gaussian, but are skewed (exhibiting an extended tail). Therefore, fitting a pure Gaussian peak to real data would not yield the best results.

To compensate for the inaccuracies involved in the conventional algorithm, the paper proposes to assess the ability of four different curve-fitting algorithms to recover the true peak shape and o/oo values of the overlapping peaks. What these four functions are may be of interest to some but aren’t relevant to this review.

Two substances exhibiting near identical o/oo values were used to experimentally generate a series of overlapping peaks. With equal sized peaks, at varying degrees of overlap (between 0% and 70%), it was observed that the conventional algorithm exhibited a depleted C13 ratio for the leading peak (-8 o/oo at maximum overlap) and an enhanced C13 ratio for the lagging peak (+8 o/oo at maximum overlap). This result was described as “unexpected”. The degree of error was proportional to the degree of overlap.

Brenna's testimony at the hearing leaned heavily on these measurements.

Application of the proposed curve fitting functions to the peaks yielded an improvement in peak area and o/oo recovery.

Similar experiments were run with a 10:1 ratio of peaks (leading peak ten times bigger than lagging peak). Using the conventional algorithm, at 40% overlap, detection of the smaller, lagging peak became problematic due to interference from leading peak’s tail. Beyond 40% overlap it was not distinguished as an individual peak.

The general trend of the leading peak being depleted and lagging peak becoming enhanced was observed for those cases where the two peaks were distinguishable.

With the peaks reversed and the leading peak being the smaller, the conventional method detected the smaller, leading peak at all degrees of overlap and it reflected a similar depletion trend as had been previously been observed.

The larger lagging peak exhibited very little error at all degrees of overlap. Curve fitting in all cases appeared to offer some advantages and generally improved the ability to recover the true o/oo of the peaks.

The curve fitting aspect is not strictly relevant to the Landis case, as it would appear that this has not used by LNDD. The conventional method and the results obtained are of more interest.

The first observation is that these results confirm Dr Brenna’s testimony of the leading peak’s C13/C12 ratio becoming depleted and the lagging peak’s C13/C12 ratio becoming enhanced. This contradicts Dr Meier-Augunstein’s testimony.

A significant factor effecting this contradiction is the effect of the m45 signal leading the m44 signal by approximately 150 ms.

Uncorrected, if the left-hand integration limit for the lagging peak is taken as the minima of the valley between the peaks and that is applied directly to both the m44 and m45 plots (not time shifted), then one would expect the C13/C12 ratio of the lagging peak to become depleted, having had a relatively larger proportion of C13 (m45) chopped off, compared to the slightly retarded C12 (m44) signal.

It would appear that performing the correction described by Brenna would resolve this issue but clearly it doesn’t. If we assume identical but scaled down peaks between the m44 and m45 plots, with no time shift (or a corrected time shift), the instantaneous 45/44 ratio would be homogeneous, presenting a constant value across the integration interval. In that case, the overlap of two peaks having identical o/oo values should have little or no impact on the measured value.

So why did overlap have such a significant effect in Brenna’s study and why were the results both contradictory to Meier-Augunstein’s opinion and described as “unexpected” by Brenna?

One possible explanation may lie in the integration time of 0.25 s.

In sampling at 0.25s intervals, how accurately will the peak maxima times on the m45 and m44 plots be identified? That’s what determines the required time shift so that they line up exactly. Remember that the required time shift is approximately 0.15s and we’re sampling at 0.25 s. A degree of error appears unavoidable. What if this error resulted in overcompensation for the time shift, resulting in a correction that had the m44 plot leading the m45 plot by some small but significant amount? This would reverse the effects described by
Dr Meier-Augenstein and result in the observations made by Dr Brenna . [un-UPDATE: incorrection removed, don't ask. ]

To this idiot, this seems the most likely reason.

In a later post, we'll look at a paper in the GDC collection, in which Brenna looks at quantization errors. These are ones where you have insufficient or mismatched sample times.

Full Post with Comments...