I will be using 8086 instruction set, since that is the most proper instruction set for beginners in my opinion.
Before getting our hands dirty with code, I will briefly explain the Von Neumann architecture, Addressing modes and registers. These are essential to be able to understand assembly language.
Main aspect of Von Neumann architecture is that the data and the instructions are stored in the memory and by looking at only the current state of the memory one is not be able to differentiate instructions from data. But CPU has a register called PC (Program Counter) which keeps the address of the next instruction to be executed. This allows CPU to differentiate between data and instructions.For example if PC never points to a memory address, that can't be an instruction. It is either not used or data that is being manipulated. This architecture will be more clear when we start seeing some code.
Addressing modes allow us to specify how "parameters" are interpreted by the CPU. For example if you want to add 5 to the register A, (You can for now think registers as variables in high-level programming languages) There are few versions of this; You can add exactly 5 to the register A, or you can add the data in the memory address 5 to the register A.
First one is called immediate addressing, while the second one is called direct addressing. (This might not seem to be direct to you since we are referring to a memory location but it is meaningful when considering the other addressing modes) I gave you only two addressing modes which are required and enough for this tutorial, but if you want to learn more about them you can always find more detailed information about them on the Internet.
Registers are very fast circuits that can store small amount of data, but have a very low access compared to memory references.
If you have been programming in a high-level language you can think of them as variables. Some of these variables are used by the CPU to get things done correctly while the others are "general purpose registers" which means you can hold any data you want in them. But the problem is they are only capable of storing 16 bits (In 8086 arch.) and we only have a few of them. (We are going to use AX,BX,CX and DX in this tutorial). Obviously 64 bits of data is not enough to code even basic programs so we will have to refer a lot to memory.
Simple instructions;
MOV AX, 0Fh
ADD AX, 0Fh
MOV BX, AX
These are simple instructions, one thing to mention is that the MOV command copies the right-value to left-location. ( MOV destination, source )
With this information we can say first instruction copies the value F to the register AX. (0xxxxxh means xxxxxx is a hex value, you better get used to hex values if you are going to program in assembly because it will make your life much easier)
Second instruction's syntax implies it's semantic which is adding the value F to AX. Now AX holds the hex value 001e and binary value 0000 0000 0001 1110. Because a hex literal corresponds to 4 bits etc.
Third instruction copies the value in the register AX to BX. Now AX and BX hold both 001e.(see? it is easier than saying 0000000000011110, this is very error-prone)
Not all instructions have two operands, for example:
INC CX
DEC DX
Can you guess what they do? Yes they increment/decrement their operands value by 1. But we can't know what CX and DX hold now. (unless we inspect flags, which we will see soon)
There are also basic logic functions:
AND AX, BX
OR AX, BX
XOR AX, BX
NOT AX
I assume you know what these are, if not you can check their truth tables on the Internet. But these instructions check the bits of their operands one by one and give corresponding bits together to XXX gates. ( XXX = AND,OR or XOR)
There is also multiplication with the instruction MUL but this is a little bit counter-intuitive, we can't do for example:
MUL only takes one operand, it expects to see the other operand on AX. So if you do
MUL 05h, the value of AX will be multiplied by 5.
You may be asking how am I going to multiply two constants then? This is how yo do it
MOV AX, constant1
MUL constant2
First you put AX the first operand and then call MUL constant2 which multiplies constant2 with AX which is constant1 and putting the result in AX.
There is also the second interesting thing, observer constant1 and constant2 can be 16 bit long which may make the product 32 bit long ! How is 32 bit product going to fit in 16-bit AX? The answer is it will not, instead CPU splits the product into two parts which are more significant part(High part) and less significant part(Low part) and puts High part to DX and Low part to AX. You can think that the DX+AX are continuous and the 32 bit product is written there.
This is the hardest part up to here, if you couldn't understand fully please read it again.
Another important thing when you are coding in assembly is LABEL's. They are not instructions nor they are interpreted by the CPU instead we use them as easy references to memory locations.
You will see codes like this:
START:
MOV AX, 05h
MOV BX, 06h
SECOND:
MUL BX
START and SECOND are labels. In this case the code does exactly the same thing with the code below. I mean EXACTLY, because unreferenced labels are simply discarded. And referenced labels are replaced by their addresses. So labels in general doesnt find a place in the generated bytecode.
MOV AX, 05h
MOV BX, 06h
MUL BX
So how does one reference a label? Meaningful way would be to use JMP but we will cover it later.
The answer is in many ways!
START:
MOV AX, SECOND
MOV BX, START
SECOND:
MUL BX
This code is perfectly fine. LABEL: INSTRUCTION means LABEL is the memory address of the following instruction. So they are just 16-bit numbers. Preprocessor runs over your code and replaces the references to labels with the memory addresses of the following instructions. So that code translates to this:
MOV AX, BEGINNING
MOV BX, BEGINNING+6
MUL BX
Here you see the labels are replaced by their "relative" addresses. +6 is there because SECOND comes after two add instructions which take up 3 bytes of memory each. So the third instruction can be inserted into BEGINNING + 6'th memory location. It is relative because Operating System can chose to place your machine code into an arbitrary memory location and give you a BEGINNING address which you can add values to it to get the desired locations.
In my compiler this code became for instance;
MOV AX, 0100h
MOV BX, 0106h
MUL BX
(so 0100 is just a "random" memory location for you)
But this usage of labels are not meaningful. We will now see JMP instruction which makes great use of labels combined with CMP instruction.
JMP takes a single operand ( generally a label in programs ) and sets PC to that value. Up to this point we wrote programs which linearly processed each instruction one by one starting from the BEGINNING. JMP changes the way this works.
MOV AX, 05h
MOV BX, 02h
JMP OVER
MOV BX, 03h
OVER:
MUL BX
Can you guess the value of BX after this code? A(2*5) or E(3*5)? Obviously A because JMP jumps to the OVER label not executing MOV BX,03h. So the third instruction never gets executed.
But this use of JMP is very static and we often need to have control over it to make meaningful things, there comes CMP instruction which means to compare.
CMP takes two operands and sets some "flags" according to it's operands. We will only consider the equality case but you can do much more with CMP and JMP instructions. There are many slightly modified JMP instructions namely
je <label> - Jump when equal
jne <label> - Jump when not equal
jz <label> - Jump when last result was zero
jg <label> - Jump when greater than
jge <label> - Jump when greater than or equal to
jl <label> - Jump when less than
jle <label> - Jump when less than or equal to
MOV AX, 09h
CMP 0Ah, AX
JE EQUAL
MOV BX, 01h ; part1
JMP FINISH ;Why do we need this unconditional JMP?
EQUAL:
MOV BX, 02h ;part2
FINISH:
INT 20h ;This is just a way to terminate the program.
This code means
if(0Ah == 09h)
execute part1
else
execute part2
If we didn't put the unconditional JMP it would mean
if(0Ah == 09h)
{execute part1}
execute part2
Observe that the part2 would get executed no matter what the result of the CMP is. This may be completely OK based on what are you trying to do, in the first code I wanted to create an if-else structure.
How does conditional JMP's know what they are going to do? They look at the flags raised by the last CMP command. But you don't have to think this in detail. Just use CMP and then a conditional JMP like JE or JNE.
No comments:
Post a Comment