序列比对工具之BLAST - 机器学习和生物信息学实验室联盟

*** Formatting options
-outfmt <String>
alignment view options:
0 = pairwise,
1 = query-anchored showing identities,
2 = query-anchored no identities,
3 = flat query-anchored, show identities,
4 = flat query-anchored, no identities,
5 = XML Blast output,
6 = tabular,
7 = tabular with comment lines,
8 = Text ASN.1,
9 = Binary ASN.1,
10 = Comma-separated values,
11 = BLAST archive format (ASN.1)
Options 6, 7, and 10 can be additionally configured to produce
a custom format specified by space delimited format specifiers.
The supported format specifiers are:
qseqid means Query Seq-id
qgi means Query GI
qacc means Query accesion
qaccver means Query accesion.version
qlen means Query sequence length
sseqid means Subject Seq-id
sallseqid means All subject Seq-id(s), separated by a ';'
sgi means Subject GI
sallgi means All subject GIs
sacc means Subject accession
saccver means Subject accession.version
sallacc means All subject accessions
slen means Subject sequence length
qstart means Start of alignment in query
qend means End of alignment in query
sstart means Start of alignment in subject
send means End of alignment in subject
qseq means Aligned part of query sequence
sseq means Aligned part of subject sequence
evalue means Expect value
bitscore means Bit score
score means Raw score
length means Alignment length
pident means Percentage of identical matches
nident means Number of identical matches
mismatch means Number of mismatches
positive means Number of positive-scoring matches
gapopen means Number of gap openings
gaps means Total number of gaps
ppos means Percentage of positive-scoring matches
frames means Query and subject frames separated by a '/'
qframe means Query frame
sframe means Subject frame
btop means Blast traceback operations (BTOP)
When not provided, the default value is:
'qseqid sseqid pident length mismatch gapopen qstart qend sstart send
evalue bitscore', which is equivalent to the keyword 'std'
Default = `0'

复制代码

4. 开始比对时，在命令行键入下列命令：
blastall -p blastn -i myRNA.fasta -d humanRNA.fasta -o myresult.blastout -a 2 -F F -T T -e 1e-10 （注意是e，不是E）
解释如下：
blastall: 这是本地化/命令行执行blast时的程序名字！(Tips:blastall直接回车就会给出你所有的参数帮助，但是英文的)
-p: p 是program的简写,program在计算机领域中是程序的意思。此参数是指定要使用何种子程序，所谓子程序，就是针对不同的需要，如核酸序列和核酸序列进行比对、蛋白质序列和蛋白质序列进行比对、假设翻译后核酸序列于蛋白质序列进行比对，选择相应的子程序: blastn 是用于核酸对核酸 blastp 是蛋白质对蛋白质序列等等，一共5个自程序。在核酸序列中搜索蛋白序列是tblastn,在蛋白序列中搜索核酸序列是blastx.
-i: i 是input的简写，意思是输入文件，就是你自己的要进行比对的序列文件(fasta格式）
-d: d是database的简写,意思是要比对的目标数据库,在例子中就是humanRNA.fasta
-o: o是output的简写，意思是结果文件名字，这个根据你自己的习惯起名字，可以带路径，(上边两个参数-i -d 也都可以带路径)

*注意以上4个参数是必须的，缺一不可，下面的参数是为了得到更好的结果自己可调的参数，如果你不加也没有关系，blastall程序本身会给一个默认值！
-a: 是指计算时要用的CPU个数，我的机器有两个CPU，所以用-a 2，这样可以并行化进行计算，提高速度，当然你的计算机就一个CPU,可以不用这个参数，系统默认值为1,就是一个CPU
-F: 是filter的简写，blastall程序中有对简单的重复序列和低复杂度的一些repeats过滤调，默认是T (注意以后的有几种参数就两个选项，T/F T就是ture,真，你可以理解为打开该功能; F就是false，假，理解为关闭该功能)
-T: 是HTML的简写，是指blast结果文件是否用HTML格式，默认是F!如果你想用IE看，我建议用-T T
-e: 是Expectation value，期望值，默认是10，我用的10-10！