1 Computational searches of biological sequences
2 Conceptos básicos Homología y otras relaciones evolutivas (paralógos, ortólogos, xenólogos) Uso preferencial de codones, CAI y expresividad Microarreglos y aproximaciones estadísticas para su análisis
3 Descripción de programas existentes BLAST (Comparación apareada de secuencias) MEME/MAST (Identificación de motivos sobre-representados)
4 Planteamiento de problemas para resolver 1. Grupo mínimo de genes para la vida 2. Predicción de operones bacterianos 3. Expresividad en unidades transcripcionales 4. Conservación de expresividad entre organismos 5. Identificación de genes transferidos horizontalmente H. pylori 6. Regulación por glucosa en E. coli
5 PAM 250 AGGIDG GHGFMG 117137 Matriz de substitución para aminoácidos
6 Unitary matrix for DNA sequences A ACGT C G T 1 1 1 1 000 0 0 0 00 0 0 0 0
7 In any case, the values obtain in the comparison are the same along the entire alignment It is well know that some residues in a protein or in a nucleotide sequence plays important roles and therefore are constrained to vary
8 These conserved regions constitutes motif which are sometimes recognized in a set of aligned sequences. SCK1_CENEL/13-36 CLKPCKDLYGPHAGAKCMNGKCKC SCKM_CENMA/13-36 CLPPCKAQFGQSAGAKCMNGKCKC SCT2_ANDAU/35-57 CASVCRRVIGVAAG-KCINGRCVC SCK3_ANDMA/13-35 CASVCRKVIGVAAG-KCINGRCVC SCK4_MESMA/35-57 CASVCRREIGVAAG-KCINGKCVC SCKK_TITSE/35-57 CYSACKKLVGKATG-KCTNGRCDC SCK2_TITDI/14-36 CVKICIDRYNTRGA-KCINGRCTC SCKP3_TITSE/7-28 CNRKCCPG-GCRSG-KCINGKCQC SCBX_MESMA/8-29 CRVKCVAM-GFSSG-KCINSKCKC SCKL_LEIQH/8-28 CQLSCRSL-GL-LG-KCIGDKCEC SCK5_ANDMA/8-28 CQLSCRSL-GL-LG-KCIGVKCEC SCK1_CENNO/36-57 CDKDCKRR-GYRSG-KCINNACKC SCK2_CENNO/8-29 CDKDCTSR-KYRSG-KCINNACKC * *. **. *
9 What is a motif in a biological sequence? Represents a conserved region of a sequence. This conservation might be due to a functional constraint. There are conserved structural domains in a family of proteins. Amino acid sequences can almost always represent such motifs. Motif identification is useful to classify and understand protein or nucleotide function.
10 Example of a protein motif. Motifs can be represented by Weight Matrices:
11 Example of a RNA motif. Motifs can be represented by Weight Matrices:
12 Example of a DNA motif. Motifs can be represented by Weight Matrices:
13 How can we obtain a Weight Matrix for a specific motif? ……. by evaluating the relative frequency of its elements in a set of aligned sequences.
14 This frequency matrix contains relevant Biological Information about your protein and can be used to obtain a: Position Specific Score Matrix PSSM
15 While PAM and Blosum matrices are used to compare two amino acids of a pair of sequences regardless of their position in the aligned sequences, a PSSM analysis uses a different matrix in which the score varies depending on the conservation of each position of the aligned sequences
16
17 Serin Protease
18 Actually, the frequencies are not used as such to score putative sites. The score assigned assigned to a piece of sequence, S, is calculated as the log-ratio of two probabilities: P(S|M), the probability to observe sequence S given the motif model M (the matrix). P(S|B), the probability to observe sequence S given the background model B (the genomic context). The score of a sequence segment is WS=log[P(S|M)/P(S|B)] Position Specific Score Matrix PSSM
19 Different programs have been developed to find motifs 1AKSJDFHLASUHERLAKSNBKAJNCLKJASHDKFJAHSEJ 2DLKTJNKHBHEASHRGHBDFASJGHBCLKUSHKLCSDHGK 3GNLKXDHKIASGCSDKJCSKHDGKJELHBHEAJFNLOIJS 4JHSLRCKJGHXBDKSLCFALSIZDNGJDFGNLCKJSDNSD 5LKSAJDHBFCKGLSHBHEAUABSXDJKFASODFHBHKAHS 6JSHGHAEKHKSDFJHKSJDFHKAJSEHRKAJHBHEAPERI 7QWHBHEACVLXMNCVKUIEHRMBDKFJAHLIDHRTRKKQP 8LICVUWJENOMNVIDFGKJERJSGFAHGSIUOPIAKHVIU 9OIEURTKSHOIUCVBSDFGUYWERKJHDFLIUHBHEAERT 10OIUWERMXCVKJHBHEAWIERUOIUVMBNAWIUEYRHASS
20 Different programs have been developed to find motifs 1AKSJDFHLASUHERLAKSNBKAJNCLKJASHDKFJAHSEJ 2DLKTJNKHBHEASHRGHBDFASJGHBCLKUSHKLCSDHGK 3GNLKXDHKIASGCSDKJCSKHDGKJELHBHEAJFNLOIJS 4JHSLRCKJGHXBDKSLCFALSIZDNGJDFGNLCKJSDNSD 5LKSAJDHBFCKGLSHBHEAUABSXDJKFASODFHBHKAHS 6JSHGHAEKHKSDFJHKSJDFHKAJSEHRKAJHBHEAPERI 7QWHBHEACVLXMNCVKUIEHRMBDKFJAHLIDHRTRKKQP 8LICVUWJENOMNVIDFGKJERJSGFAHGSIUOPIAKHVIU 9OIEURTKSHOIUCVBSDFGUYWERKJHDFLIUHBHEAERT 10OIUWERMXCVKJHBHEAWIERUOIUVMBNAWIUEYRHASS …..if the alignment is not an option?
21 How do they work? A) Counting all the “words” of certain length and evaluating the more frequent and statistically significant. B) In a aleatory fashion, taking fragments chosen randomly and evaluating if these fragments manage to generate a conserved representative motif Gibbs sampler algorithm) ( Gibbs sampler algorithm)
22 Gibbs sampler algorithm Multiple Local Alignment (MLA)
23 We mark a sequence into the motif site (occurrence), which is described by a probability-positional matrix q(i,r), and the background, which is described by background symbol probabilities f(i). r is a nucleotide (a residue); r {A,T,G,C} i is a position in the site, i=1..s, s is the motif length Positional-Probabilistic Model (PPM) and background
24 What is a motif Two probabilistic models, foreground (the motif) and background, are formulated. We classify (mark) all the input sequences into these two models-obtained parts.
25 A Gibbs sampling step Motif and background bases counters are computed from all the sequence fragments except the current one. The probability distribution of the new site position or its absence in the current sequence is derived from the statistical models and the current sequence content. A new site location is sampled from the distribution. Statistical models for the background and for the motif are formed using the counters. The current sequence
26 …..if the alignment is not an option? Gibbs sampler algorithm
27 Motif site (occurrence), which is described by a probability-positional matrix q(i,r) background, which is described by background symbol probabilities f(i). Two probabilistic models are formulated: the foreground model (the motif) and the background model
28 Gibbs sampler algorithm A probability distribution (where the foreground and background models are different) can be evaluated
29 A complete statistical description of the method is not in the scope of this talk
30 One of the sequences, chosen randomly, is removed from the alignment. The main idea of the method ….. A probability distribution profile is evaluated
31 and replaced by new sequence searched with the previous motif profile A new probability distribution profile is evaluated again The main idea of the method …..
32 After several cycles, the method tends to identify a significant motif
33 GLAM2 is a software package for finding motifs in sequences, typically amino- acid or nucleotide sequences. The main innovation of GLAM2 is that it allows insertions and deletions in motifs. The package includes these programs: * glam2 - for discovering motifs shared by a set of sequences. * glam2scan - for finding matches, in a sequence database, to a motif discovered by glam2. * glam2format - for converting glam2 motifs to standard alignment formats. * glam2mask - for masking glam2 motifs out of sequences, so that weaker motifs can be found. * purge - for removing highly similar members of a set of sequences. http://bioinformatics.org.au/glam2/doc/
34 http://meme.sdsc.edu/meme4/cgi-bin/glam2.cgi
35 Basic usage Running glam2 without any arguments gives a usage message: Usage: glam2 [options] alphabet my_seqs.fa Main alphabets: p = proteins, n = nucleotides Main options (default settings): -h: show all options and their default settings -o: output file (stdout) -r: number of alignment runs (10) -n: end each run after this many iterations without improvement (10000) -2: examine both strands forward and reverse complement --z: minimum number of sequences in the alignment (2) --a: minimum number of aligned columns (2) --b: maximum number of aligned columns (50) --w: initial number of aligned columns (20) -The main input to glam2 is a file of sequences in FASTA format: >MyFirstSequence GHYWVVCTGGGACH >My2ndSequence LLIGGPWVWWADDDF (etc.) You need to tell glam2 which alphabet to use: glam2 p my_prots.fa glam2 n my_nucs.fa Use -o to write the output to a file rather than to the screen: glam2 -o my_prots.glam2 p my_prots.fa
36 How it works To use glam2 starts from a random alignment, and makes many small, random changes to it, which are designed to find high-scoring alignments in the long run. The longer you let it run, the more likely it is to find a maximal-scoring alignment. To check that a reproducible, high-scoring motif has been found, the whole procedure is run several (e.g. 10) times from different starting alignments. If all runs produce identical alignments, we have maximum confidence that this is the optimal motif. If a few of the runs produce different, lower-scoring motifs, we still have high confidence. If all the runs produce completely different alignments, we have low confidence, and the run-length needs to be increased.
37 MEME MEME: Multiple Expectation maximization for Motif Elicitation MAIN DIFFERENCES Try many different initial segments in order to get one that converges to an optimum. Try different window analysis sizes. In order to generate a motif with gaps, more than one motif can be generated.
38 Motif discovery from unaligned sequences Optimal for Genomic or Protein sequences Especially if you do not know the size of the motif Identifies profile motifs Simultaneously analyze Multiple motifs for any input Flexible model of motif presence Motif can be absent in some sequences Motif can appear several times in one sequence A very useful program to discover Patterns MEME MEME: Multiple Expectation maximization for Motif Elicitation
39 The input to MEME contains the following fields. Sequences Protein or DNA sequences (Do not merge) in fasta format. Notice that sequence names must not be repeated Valid examples are >seq1 GDIFYPGYCPDVKPVNDFDLSAFAGAWHEIAK >seq2 GDMFCPGYCPDVKPVGDFDLSAFAGAWHELAK >GdhA Glutamate dehydrogenase from Escherichia coli RDMFCPGGCPDVKPVGDHDLSAFKGAWHELAL
40 Motif distribution One per sequence (oops) Zero or one per sequence (zoops) Any number of repetitions (tcm) Number of motifs [optional] The program will stop the analysis after this number of motifs is found. Number of sites (Minimum or Maximum sites) (
41 Text output format By default, MEME output is in hypertext (HTML) format. Shuffle letters in input sequences Useful for further statistical analysis Look for palindromes only Average the letter frequencies in corresponding motif columns together MEME input. continued….
42 http://meme.sdsc.edu/meme/meme.html
43 MEME Output Motif length # of motifs found Expectation value “Position-Specific Probability Matrix Information content Consensus
44 Sequence names Strand (reverse or complement) Position in sequence Statistical significance Motif within sequence MEME Output
45 Overall Statistical significance Sequence length Motif in complement strand MEME Output
46 Searches for motifs (one or more) in sequence databases: Like BLAST but motifs are used as input Similar to the matrices obtained by iteration in PSI- BLAST Profile defines statistical significance of a match Multiple motif matches per sequence Combined E value for all motifs MEME uses MAST to summarize results: Each MEME result is accompanied by the MAST result for searching the discovered motifs on the given sequences. MAST
47 Email address Database (like BLAST) Motif file (e.g. MEME output) Consider matched sequence length E value threshold MAST input
48 Matched accession Match E value Length of sequence Link to GenBank MAST output
49 Motif diagram MAST output
50 Position of each instance P value of instance Matched parts of sequence Motif ‘consensus’ Motif and orientation MAST output
51 Affymetrix GeneChip arrays High density oligonucleotide array technology as developed by Affymetrix www.affymetrix.com Overview images courtesy of Affymetrix unless otherwise specified
52 Probes and Probesets
53 Sample Preparation
54 Hybridization to the Chip
55 The Chip is Scanned
56 The task of analysing a gene expression experiment for differential genes falls into the following steps: (1) Ranking: genes are ranked according to their evidence of differential expression. (2) Assigning significance: a statistical significance is being assigned to each gene. (3) Cut-off value: to arrive at a limited number of differentially expressed genes a cut-off value for the statistical significance needs to be determined. DIFFERENTIAL EXPRESSION Motivation
57 Planteamiento de problemas para resolver 1. Grupo mínimo de genes para la vida 2. Predicción de operones bacterianos 3. Expresividad en unidades transcripcionales 4. Conservación de expresividad entre organismos 5. Identificación de genes transferidos horizontalmente H. pylori 6. Regulación por glucosa en E. coli
58 Datos de microarreglos sobre la reglulación por glucosa en E. coli IDB_numbercrp1 crp2 wt1 wt2 wtg1 wtg2wtg3 señaldetseñaldetseñaldetseñaldetseñaldetseñaldet 1aas_b2836_at279.2P293.8P360.8P332.2P475.6P374.4P 2aat_b0885_at344.9A332.5A110.7A256.8A237.1A35.2A 3abc_b0199_at697M682.4A530.5A532A747.9M882.8A 4abrB_b0715_at537.3P518.3P446.3A329.6M357.9P406.7A 5accA_b0185_at1735.3P1083.1P1821.7P1798.6P2416.7P1794.5P 6accB_b3255_at2314.2P1811.4P1598P1835.2P3182.8P2893.6P 7accC_b3256_at2486.2P1845.6P1165.7P1146P2861.6P2883.3P 8accD_b2316_at1770.3P1468.1P1592.1P1808.5P2257.6P2873.7P 9aceA_b4015_at632.6P790.2P3005.5P2489P474.1P521.8P 10aceB_b4014_at282.7P394.4P1610.5P1254.1P274.8A316.6A 11aceE_b0114_at5060.1P3979.2P2093.5P1665.9P7943.2P8411.5P 12aceF_b0115_at3981.4P2287.3P1998.1P1425.8P5515.2P4993.9P 13aceK_b4016_at386.8P331P534.9P442.9P395.4P395.3P 14ackA_b2296_at2624.8P2084.2P1225.1P1273P3624.6P2990.2P 15acnA_b1276_at488P326.2P701.9P773.2P368.7A460.8A 16acnB_b0118_at2417.2P2185.7P3693P2864.5P1401.4P1786P P-Confiable, A-Dudoso, M-no confiable
59 Enfoques de estudio Tecnología de microarreglos Se pueden complementar con otros métodos (RT-PCR, por ejemplo)
60 Outliers iteration method 1.Se obtienen los promedios de las diferentes repeticiones 2.Se comparan los resultados de los experimentos de estudio con los del experimento control. En nuestro ejemplo: a) WTglucosa/WTb) CRP/WT 3. Se obtienen la media y desviación estandar de los datos.
61 4.- Para identificar a los genes con mayor significancia asumimos un corte arbitrario de 2 desviaciones estardar relativos a la media. Outliers iteration method Genes inducidos Genes reprimidos
62 5.- Se vuelve a calcular la media y desviación estandar de los datos que no fueron considerados como estadísticamente importantes. Outliers iteration method
63 6.- Nuevamente se considera que los genes por arriba o debajo de dos desviaciones 2 desviaciones estardar relativos a la media son estadísticamente relevantes Outliers iteration method Genes inducidos Genes reprimidos
64 Outliers iteration method Se itera hasta ya no encontrar valores por afuera de las 2 DS. Genes inducidos Genes reprimidos Genes inducidos Genes reprimidos
65 Metodo de ordenamiento (Rank) Destaca los genes más expresados. Por lo tanto, no importa el valor en términos absolutos si no su valor relativo a los demás valores del arreglo Los genes en principio deben de conservar su posición dada la naturaleza de su expresión.
66 Metodo de ordenamiento (Rank) GeneValor A0.3 B0.6 C1.4 D3.2 E1.1 F4.3 G0.9 H0.8 I0.4 J0.7 PosGeneValor 1F4.3 2D3.2 3C1.4 4E1.1 5G0.9 6H0.8 7J0.7 8B0.6 9I0.4 10A0.3
67 Metodo de ordenamiento (Rank) PosGeneValor 1F4.3 2D3.2 3C1.4 4E1.1 5G0.9 6H0.8 7J0.7 8B0.6 9I0.4 10A0.3 ¿Cuál es la probabilidad de quedar en primer lugar? P-primero = 1/n = 1/10 = 0.1 ¿Cuál es la probabilidad de quedar en primer o segundo lugar? p= p-primero + p-segundo = 0.2 ¿Cuál es la probabilidad de quedar en segundo lugar? P-segundo = 1/n = 1/10 = 0.1
68 Metodo de ordenamiento (Rank) PosGeneValor 1F4.3 2D3.2 3C1.4 4E1.1 5G0.9 6H0.8 7J0.7 8B0.6 9I0.4 10A0.3 PosGeneValorProb 1F4.30.1 2D3.20.2 3C1.40.3 4E1.10.4 5G0.90.5 6H0.80.6 7J0.7 8B0.60.8 9I0.40.9 10A0.31.0
69 Metodo de ordenamiento (Rank) ¿Cómo consideramos la probabilidad si existen repeticiones del experimento? Media geométrica 1er exp -> 3er lugar -> 0.3 2do exp -> 1er lugar -> 0.1 3er exp -> 2do lugar -> 0.2 P = (0.3 x 0.1 x 0.2) 1/3 P = (0.06) 1/3 P = 0.39
70 Problemas a resolver sobre análisis de microarreglos 1. ¿Cuáles son los genes reprimidos e inducidos cuando E. coli es crecida en glucosa? 2. ¿Cuáles son los genes reprimidos e inducidos en la mutante CRP de E. coli? 3. ¿Cuál es la coincidencia entre los genes de los puntos 1 y 2? 4. ¿Qué tanto coinciden los resultados del análisis de Outliers y Rank? 5. Identifique el posible elemento de regulación común a estos genes utilizando Glim2 y MEME
71