I'm trying to extract information from certain rows in this big dataframe.
When I do use slicing to subset the table (e.g. blast_output_scored.iloc[10:11,:]), The output looks like this:
| qseqid | sseqid | %_identity | alignment_length | mismatch | gapopen | qstart | qend | sstart | send | evalue | bitscore | subject_strand | line_in_og_BLAST | Needle_score |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| IDgene.1 | 1 | 100.0 | 1073 | 0 | 0 | 1 | 1073 | 7704 | 6632 | 0.0 | 1982.0 | minus | 10 | 5360.0 |
When I go to check the number of rows in this table, I get the correct number with slicing
len(blast_output_scored.iloc[10:11,:].index)
#Output
1
However when I use indexing (blast_output_scored.iloc[10,:]), the output looks completely different, even if the index is the same range as the slice.
sseqid 1
%_identity 100.0
alignment_length 1073
mismatch 0
gapopen 0
qstart 1
qend 1073
sstart 7704
send 6632
evalue 0.0
bitscore 1982.0
subject_strand minus
line_in_og_BLAST 10
Needle_score 5360.0
Name: IDgene.1, dtype: object
The number of rows in the table now doesn't seems to change to the number of columns - the first column (the first column is also set for indexing rows by names)
len(blast_output_scored.iloc[10,:].index)
#Output
14
My biggest problem is that I'm using the column names to index and I have to check which names subset tables to a length of 1, so I can't just use the splicing method to bypass this.
e.g. blast_output_scored.loc["IDgene.1"] outputs
sseqid 1
%_identity 100.0
alignment_length 1073
mismatch 0
gapopen 0
qstart 1
qend 1073
sstart 7704
send 6632
evalue 0.0
bitscore 1982.0
subject_strand minus
line_in_og_BLAST 10
Needle_score 5360.0
Name: IDgene.1, dtype: object
and will say I have 14 rows when I should only have 1.
Is there any way to ensure the output looks like the slicing output in pandas?