Académique Documents
Professionnel Documents
Culture Documents
It is also advisable to have on-chip debugging facility (through JTAG port) and debug
software.
6. Availability of the microcontroller chips is also another determining factor.
Several microcontrollers available in the market. The following are some of the important
ones 68HC11, 8051, ARM, Atmel (AVR8, AVR32) , Freescale (CF, S08), Hitachi (H8, SuperH), MIPS, NEC, PIC, PowerPC, TI MSP430, Toshiba TLCS-870, Zilog (eZ8, eZ80). In the
following, we look into the ARM processor architecture that has many nice features supporting
embedded computation, and is in particular, low power.
1.1
ARM microcontroller
ARM is a 32-bit RISC (Reduced Instruction Set Computer) processor architecture developed
by the ARM Corporation. It was previously known as Advanced RISC Machine, and prior
to that Acron RISC Machine. The ARM architecture is licensed to companies that want to
manufacture ARM based CPUs or system-on-a-chip products. This enables the licensee to
develop their own processors compliant with the ARM instruction set architecture. ARM processors possess a unique combination of features that makes ARM the most popular embedded
architecture today.
1. ARM cores are very simple, compared to other general-purpose processors available in
the market. This implies that the ARM processors will need relatively lesser number of
transistors, leaving enough space on the die to realize other functionalities on the silicon.
2. The instruction set architecture and the pipeline design of ARM are aimed at minimizing
the energy consumption a critical requirement in mobile embedded systems. In spite
of being a 32-bit microcontroller, it is capable of running 16-bit instruction set, known
as THUMB. This helps to achieve greater code density and enhanced power saving.
3. While being small and low-power, ARM processors provide high performance.
4. ARM architecture is highly modular the only mandatory component of an ARM processor is the integer pipeline. All other components including caches, memory management
unit (MMU), floating point and other co-processors are optional. This gives a lot of
flexibility in building application specific ARM-based processors.
5. To assist the developer, the ARM core has a built-in JTAG debug port and on-chip
embedded ICE (In-Circuit Emulator) that allows programs to be downloaded and
fully debugged in-system.
1.2
A brief history
The first ARM processor was designed by the Acron Computers Limited of Cambridge, England
between 1983 and 1985. While looking for a new processor for the next generation desktop,
they found that the existing commercial microprocessors were not suitable for their purpose,
mainly because of the following two reasons.
These processors were slower than the existing memory parts.
Complex instruction set including instructions requiring hundreds of cycles to execute,
leading to high interrupt latencies.
Thus, the requirement of designing a new processor was felt. However, designing a complex
processor used to take many years even for large companies with expertise available for processor design. The solution was found in the Berkeley RISC 1 project which had established
that it was possible to build a very simple processor with performance comparable to the most
advanced CISC processors of the time. Acron designed their first 26-bit Acron RISC Machine
(ARM) processor in 1985, based upon the Berkeley project. It used even lesser than 25,000
transistors and still performed similar to (or better than) Intel 80286 processor that came
out at about the same time. This architecture has later been referred to as ARM version
1 architecture. It was followed by the second processor in 1987. This ARM version 2 had
a coprocessor support. It was extended with on-chip cache in ARM 3 processor. The third
version of ARM architecture, developed in 1992 had features like 32-bit addressing, support
for Memory Management Unit (MMU) and 64-bit multiply-accumulate instructions. It was
implemented in ARM 6 and ARM 7 cores. Prior to this, in 1990, Apple took a decision to use
ARM processor in their Newton PDA. A joint venture called ARM (Advanced RISC Machines)
was launched. ARM entered into the embedded market with the release of these processors
and the Apple Newton PDA in 1992. The 4th generation ARM processor came out in 1996
with the special feature of THUMB 16-bit compressed instruction set. Though THUMB
is slightly less efficient compared to the regular 32-bit ARM instruction set, it takes 40% less
space. The most prominent representative of the 4th generation ARM is the ARM7TDMI core,
which is till now the most popular ARM product. It has been used in most Apple iPod players,
including the video iPod. Another popular implementation of ARMv4 core is the Intel StrongARM processor. The 5th generation of ARM architecture introduced in 1999 have digital
signal processing capability and Java byte code extensions to the ARM instruction set. Intel
XScale processor is the most popular implementation of the 5th generation ARM core. It is
used in a number of embedded devices, network processors, smart phones and PDAs. ARMv6
architecture anounced in 2001 features improvements in many areas covering the memory sys-
tem, improved exception handling and better support for multiprocessing environments. It
also includes media instructions to support Single Instruction Multiple Data (SIMD) software
execution. THUMB-2, an improved THUMB instruction set defining a new set of 32-bit instructions that execute along-side 16-bit instructions in THUMB state, was introduced. It
provides better support for two separate address spaces, such that code executing in the nonsecure world cannot gain access to any address space marked as secured. The protection
provided by the technology is necessary for consumer privacy and extending a range of services
such as, mobile banking and multimedia entertainment, to widespread consumer adoption and
use. The next generation ARMv7 cores have been introduced in 2005. It comes with three
different processor profiles. The A profile is for sophisticated virtual memory based OS and
user applications. The R profile is for real-time systems, and the M profile is optimized
for microcontrollers and low-cost applications. The ARMv7A architecture has the option of
NEON technology designed to address the next generation high performance, media intense,
low-power mobile hand-held devices. It is a 64/128-bit hybrid SIMD architecture developed
by ARM to accelerate the performance of multimedia and signal processing applications. The
Vector Floating Point (VFP) coprocessor support is also an architectural option. It supports
single and double precision floating point arithmetic, and is fully IEEE 754 compliant with
suitable software library. Table 1.1 summerizes the discussion.
Version
v1
v2
v3
v4
Year
1985
1987
1992
1996
v5
v6
v7
1999
2001
2005
1.3
Structure of ARM7
Fig 1.1 shows a block diagram of ARM7 processor. The major components of ARM7 processor
are described in the following.
1. Instruction Pipeline and Read Data Register: It gets the content of memory location
pointed to by the address bus lines A[31 : 0], of Address Register. The external 32-bit
data-in lines DAT A[31 : 0] puts the content into this register.
2. Instruction Decoder and Control Logic: It has a number of control inputs determining the
operation policy of the processor. Also, it outputs a number of control signals useful for
interfacing the processor with other peripherals. The various control signals are explained
later.
3. Address Register: It holds the address of the next instruction/data to be fetched. Address
bus A[31 : 0] originates from it. The input signal ALE determines the time upto which
the registers content will remain available on the A[31 : 0] lines. Content is available as
long as ALE remains low.
4. Address Incrementer: It increments the Address Registers value by an appropriate
amount to point to the next instruction/data.
5. Register Bank: It contains 31, 32-bit registers accessible in different modes of operation
of the processor (detailed later). It also contains 6 status registers, each of size 32 bits.
6. Booths Multiplier: It is used in the multiplication instructions.
7. Barrel Shifter: One of the opearnds of data processing instructions can be shifted by a
few bit positions. The barrel shifter located at the input of ALU performs this function.
8. ALU: A 32-bit ALU performs the arithmetic and logic functions.
9. Write Data Register: It holds the value to be written into the memory. The 32-bit value
is available in the DOU T [31 : 0]. The associated signals DBE and nENOUT have been
elaborated later.
1.3.1
We will next elaborate the control and status signals used in ARM7 processor. Fig 1.2 shows
a functional diagram of ARM7 CPU. The signals can be grouped as per their functionality.
That way, the signals can be grouped as,
Processor mode: includes the signals nM [4 : 0].
Memory interface: A[31 : 0], DAT A[31 : 0], DOU T [31 : 0], nEN OU T , nM REQ, SEQ,
nRW , nBW , LOCK.
Memory management interface: nT RAN S, ABORT .
ALE A[31:0]
PC Bus
Address
Incrementer
ALU Bus
B Bus
Register Bank
(31 X 32bit registers)
(6 status registers)
Incrementer Bus
Address Register
A Bus
Instruction
Decoder
&
Control
Logic
Booths
Multiplier
Barrel
Shifter
32bit ALU
Instruction Pipeline
& Read Data Register
DATA[31:0]
nEXEC
DATA32
BIGEND
PROG32
MCLK
nWAIT
nRW
nBW
nIRQ
nFIQ
nRESET
ABORT
nOPC
nTRANS
nMREQ
SEQ
LOCK
nCPI
CPA
CPB
nM[4:0]
Memory interface signals A[31 : 0], DAT A[31 : 0], DOUT [31 : 0], nENOUT ,
nMREQ, SEQ, nRW , nBW , LOCK
A[31 : 0] are the address lines constituting the processor address bus. If ALE is high, addresses
become valid during the phase 2 of the previous instruction cycle. This address is used in phase
1 of the referenced cycle. The stable period may be controlled by ALE as discussed later.
DAT A[31 : 0] is the input data bus. During read cycles (identified by nRW = 0), input must
be valid before the end of phase 2 of the transfer cycle. DOU T [31 : 0] is the output data bus.
During write cycles (nRW = 1), outout data becomes valid during phase 1 and remain so for
the entire phase 2 of the transfer cycle. nEN OU T is a status signal which is activated (made
low) by the processor when DOU T contains a valid data to be written into the memory. This
nEN OU T signal can be utilized to create a bidirectional bus with DAT A for the memory.
nM REQ is a status signal indicating (when low) that the processor requires memory access
during the following cycle. The signal becomes valid during phase 1, remaining valid through
phase 2 of the cycle preceding that to which it refers.SEQ active-high signal indicates that
the address used in the following cycle is either the same as the last memory address, or is
4 greater (ie. the next word address). It becomes valid during phase 1 and remains valid
throughout phase 2 of the cycle before the one to which it refers. The two signals nM REQ
and SEQ together indicate burst activity one cycle advance. nRW is another status signal.
For a read cycle, the signal is low, for a write cycle, it is high. nBW is high for a word transfer
and low for a byte transfer.The signal LOCK is used for locked memory access. When LOCK
is high, the memory controller should not allow any other device to access memory till LOCK
becomes low. It is used in particular in swap instruction.
Interrupts nIRQ, nF IQ
nF IQ is an asynchronous interrupt to the processor. The signal is low-level sensitive and must
be held low till an action is received from the processor. This is responded the fastest. On the
other hand, nIRQ is of lower priority.
1.4
ARM Pipeline
One important feature of ARM architecture is its pipelined organization. There are various
versions of this pipeline used in different ARM architectures. These are,
3-stage pipeline (ARM7TDMI and earlier).
5-stage pipeline (ARMS, ARM9TDMI).
6-stage pipeline (ARM10TDMI).
8-stage pipeline (ARM11).
1.4.1
3-stage pipeline
The 3-stage pipeline is the classical fetch-decode-execute pipeline. The first pipeline stage reads
an instruction from memory and increments the value in the instruction address register. The
value is also stored in the program counter (PC). The next stage decodes the instruction and
prepares control signals required to execute it on. The third stage does all the actual works:
reading operands from the register file, performing ALU operations by fetching or writing to
the memory the temporary data (if necessary), and finally writing back the modified register
values. In the case of data processing instructions, the result from ALU is directly written into
the register file, thus completing the execution stage of the instruction in one cycle, while for
10
MCLK
Clocks
nM[4:0]
nWAIT
Processor
Mode
A[31:0]
PROG32
Configuration
DATA32
BIGEND
DATA[31:0]
DOUT[31:0]
nEXEC
nENOUT
nMREQ
nIRQ
Interrupts
nFIQ
nRESET
ARM7
Memory
Interface
SEQ
nRW
nBW
LOCK
nTRANS
Bus
Controls
ALE
ABORT
DBE
Memory
Management
Interface
nOPC
Power
VDD
nCPI
CPA
VSS
CPB
Coprocessor
Interface
11
load/store type of instructions, the address computed by the ALU is placed on the address bus
and the actual memory access is performed during the second cycle of the execution stage.
1.4.2
5-stage pipeline
The problem with 3-stage pipeline is the pipeline stall caused by every data transfer instruction
the next instruction cannot be fetched while memory is being read/written. To circumvent
the problem, in ARM9TDMI and later architectures, instruction and data memory have been
separated. To make the pipeline balanced the following measures have been taken.
The register read step is moved to the decode stage.
Execute stage is split into three stages. The first stage performs arithmetic computations,
the second stage performs memory access, while the third stage writes the result back to
the register file. It may be noted that the second stage will remain idle when executing
data processing instructions.
This modification balances the pipeline, reducing the CPI (average number of Clocks Per
Instruction). However, there is a new complicacy we need to forward data between pipeline
stages to resolve data dependencies between the stages without stalling the pipeline.
1.4.3
6-stage pipeline
In ARM10 core, the instruction decode stage is split into two pipeline stages the decode
stage and the register stage. This creates a 6-stage pipeline. While the decode stage performs
the decoding operation, register stage reads the register to be used. The major advancement
introduced are in the width of the instruction and data buses, both of which are made 64-bit.
Thus, fetch stage can fetch two instructions simultaneously. A static branch predictor module
has been introduced. A separate adder has been introduced in the execution unit to take care
of multiply-accummulate imstructions.
1.4.4
8-stage pipeline
Two major changes have been introduced in ARM11 cores creating an eight-stage pipeline.
Shift operation has been separated into a separate pipeline stage.
12
Both instructions and data accesses are distributed across two pipeline stages.
The execution unit is split into three different pipelines that can operate concurrently and
commit instructions out-of-order also.
1.5
13
the condition stated must be met. This goes a long way to eliminate small branches in
the program code and thus eliminating stalls in the pipeline.
All these features have resulted in high performance, low code size, low power consumption,
and low silicon area in ARM.
1.5.1
Registers
The ARM ISA has 16 general-purpose registers, R0 R15, in the user mode. Out of this, register R15 is the program counter which may also be manipulated as a general-purpose register.
Register R13 and R14 also have special functions; R13 is used as the stack pointer, though
this has only been defined as a programming convention. Unusually, the ARM instruction set
does not have PUSH and POP instructions, so stack handling is done via a set of instructions
that allow loading and storing multiple registers in a single operation. R14 has special significance and is called the link register. When a procedure call is made, the return address is
automatically placed into this register (and not in the stack, as usually done in other processors). A return from the procedure can thus be implemented by moving the content of R14
to R15. Another important register, the current program status register (CPSR) contains four
1-bit condition flags (namely, negative, zero, carry, and overflow) and four fields representing
the execution state of the processor. The I and F flags of CPSR enable normal and fast
interrupts respectively. The T field is used to switch between ARM and THUMB instruction
sets. The mode field selects one of the six execution modes as follows.
1. User mode is used to run the application code. Once in user mode, the CPSR cannot be
written to. Mode can only be changed when an exception is generated.
2. Fast interrupt processing mode (FIQ) supports high speed interrupt handling. Generally
it is used for a single critical interrupt source in a system.
3. Normal interrupt processing mode (IRQ) supports all other interrupt sources in a system.
4. Supervisor mode (SVC) is entered when the processor encounters a software interrupt
instruction. These are standard ways to invoke operating system services. Upon reset,
ARM enters into this mode.
5. Undefined instruction mode (UNDEF) is entered if the fetched opcode is not an ARM
instruction or a coprocessor instruction.
6. Abort mode is entered in response to memory fault, for example, an instruction or data
fetched from an invalid memory region.
14
The user registers R0 to R7 are common to all operating modes. However, FIQ mode has
its own R8 to R14 registers that replace the user registers when FIQ mode is entered. Similarly,
each of the other modes have their own R13 and R14 registers so that each mode has its own
stack pointer and link register. The CPSR is also common to all the modes. However, in each
of the exception modes, an additional register the saved program status register (SPSR) is
added. SPSR registers store a copy of the value of the CPSR register before an exception was
raised. Fig 1.3(a) shows all the user accessible registers while Fig 1.3(b) shows the structure
of the CPSR.
User
R0
R1
R2
R3
R4
R5
R6
R7
R8
R9
R10
R11
R12
R13
R14
R15 (PC)
FIQ
R0
R1
R2
R3
R4
R5
R6
R7_fiq
R8_fiq
R9_fiq
R10_fiq
R11_fiq
R12_fiq
R13_fiq
R14_fiq
R15 (PC)
Supervisor
R0
R1
R2
R3
R4
R5
R6
R7
R8
R9
R10
R11
R12
R13_svc
R14_svc
R15 (PC)
Abort
R0
R1
R2
R3
R4
R5
R6
R7
R8
R9
R10
R11
R12
R13_abt
R14_abt
R15 (PC)
IRQ
R0
R1
R2
R3
R4
R5
R6
R7
R8
R9
R10
R11
R12
R13_irq
R14_irq
R15 (PC)
Undefined
R0
R1
R2
R3
R4
R5
R6
R7
R8
R9
R10
R11
R12
R13_und
R14_und
R15 (PC)
CPSR
CPSR
SPSR_fiq
CPSR
SPSR_svc
CPSR
SPSR_abt
CPSR
SPSR_irq
CPSR
SPSR_und
(a)
31 30 29 28 27
N Z C V
8 7 6 5 4
Unused
N: negative
Z: zero
C: carry
V: overflow
F T
0
Mode
I: IRQ enable
F: FRQ enable
T: THUMB instruction set
Mode: 6 operating modes
(b)
Figure 1.3: (a) ARM registers in different modes (b) CPSR register content
15
Data Types
The ARM instruction set supports six different data types, namely, 8-bit signed and unsigned,
16-bit signed and unsigned, 32-bit signed and unsigned. The ARM processor instruction set
has been designed to support these data types in little- or big-endian format. However, most
of the ARM silicon implementations use the little-endian format.
In the following we give a brief overview of different types of ARM instructions. ARM has
got two instruction sets:
ARM:
Standard 32-bit instruction set
It consists of the following types of instructions as shown in Fig 1.4.
Data processing
Data transfer
Block transfer
Branching
Multiply
Conditional
Software interrupts
THUMB:
16-bit compressed form
Code density better than most CISC processors
Dynamic decompression in pipeline
1.5.2
The ARM architecture provides a range of arithmetic operations, such as, addition, subtraction,
multiplication etc., and a set of bit-wise logical operations. All these instructions take two 32bit operands and return a 32-bit result. The multiplication instruction can return a 32- or
64-bit value. All these operands and results can be specified independently. Out of the three,
the first operand and the result must be registers, while the second operand can be either a
register or an immediate value. If the second operand is a register, it can be shifted or rotated
16
R1 = R2 + R3.
R1 = R2 + (R3 4).
R1 = R2 + (R3 4) and set condition code flags.
1.5.3
17
ARM supports two types of data transfer instrcutions: single register transfers and multiple
register transfers. Single register transfer instructions can be used to transfer 1, 2, or 4-bytes of
data between a register and a memory location. On the other hand, multiple register load/store
operations can be carried out via multiple register transfer instructions. The addressing mode
to be used for multiple register transfer is base-plus-offset addressing. Value in the base register
is added to the offset stored in a register or passed as an immediate value to form the memory
address. An auto-indexed addressing mode can also be used which writes the value of the
base register incremented by the offset back to the base register. It helps in accessing the
next memory location in the next instruction without wasting an additional instruction to
increment the register.Two different auto-indexed addressing modes are supported the preindexed mode uses the new computed address for the load/store operation, and then updates
the base register with the new computed value. The post-indexed mode uses the unmodified
base register for the transfer and then updates the base register. Multiple register transfer
instructions also support auto-indexed addressing. Multiple register transfer instructions are
particularly useful while entering or exiting a procedure to pass parameters or return values.
The following points may be noted regarding the offset and the indexing.
18
LDR
LDR
LDR
LDR
R0,
R0,
R0,
R0,
[R8]
[R1, -R2]
[R1, #4]
[R1, #4]!
pointed
pointed
pointed
pointed
to
to
to
to
by
by
by
by
R8 into R0
R1 R2 into R0
R1 + 4 into R0
R1 + 4 into R0,
LDR
Load word
STR
Store word
LDRH
STRH
LDRSH
STRSH
LDRB
Load byte
STRB
Store byte
LDRSB
STRSB
As discussed earlier, ARM supports both little-endian and big-endian formats for data
access. In the little-endian format, the least significant byte of a word is stored in bits 0-7 of
an addressed word, whereas, in the big-endian format, the least significant byte is stored in
bits 24-31. It may be noted that this has got significance only when the data is stored as words
and then accessed as bytes or halfwords. Fig 1.5 gives an example of the same.
R0 = 0x10203040
31
24 23
10
8 7
16 15
20
30
40
Memory
31
Memory
24 23
40
30
8 7
16 15
20
20
8 7
16 15
00
00
10
8 7
16 15
30
40
24 23
00
24 23
10
R1 = 0x200
10
Bigendian
31
31
Littleendian
31
24 23
00
00
R2=0x40
R2= 0x10
8 7
16 15
00
40
19
20
Old SP
Old SP
SP
r5
r4
r3
r1
r0
r5
r4
r3
r1
r0
r5
r4
r3
r1
r0
Old SP
SP
Old SP
r5
r4
r3
r1
r0
SP
1.5.4
21
Multiplication instructions
Multiply
Multiply accumulate
Unsigned multiply
Unsigned multiply accumulate
Signed multiply
Signed multiply accumulate
32-bit result
32-bit result
64-bit result
b 64-bit result
64-bit result
64-bit result
For example,
MUL R0, R1, R2
MULA R0, R1, R2, R3
R0 gets R1 R2
R0 gets R1 R2 + R3
22
In some other variants containing extended multiplication hardware, the following enhancements have been performed.
An 8-bit Booths algorithm is used, which makes the multiplication faster (maximum
standard instruction is now 5 cycles).
Early termination method is improved in the sense that the multiplication completes
when all remaining bit sets contain either all zeros, or all ones.
64-bit results can be produced from 32-bit operands, this provides higher accuracy. A
pair of registers is used to hold the result.
1.5.5
Software Interrupt
The software interrupt (SWI) instruction forces the CPU into supervisor mode. Its format is,
SWI #n
The execution of the instruction causes an exception trap to the SWI hardware vector (forcing
the mode change and an associated state saving), thus causing the SWI exception handler to
be called. The handler can now analyze the value of n to determine which action to perform.
It should be noted that the processor completely ignores n, and it is the software interrupt
handler that may analyze the value of n to perform the desired job. This is very suitable for
an operating system to implement a set of privileged operations which applications running in
user mode can request. Such requests are commonly known as system calls. The value of n is
24-bits, thus providing facility for a maximum of 224 calls. Unlike many other processors (such
as, Intel processors), the hardware does not attempt to distinguish between these 224 cases. If
separate addresses are to be maintained for them, a total of 224 4 (address bus is 32 bits =
4 bytes) = 226 bytes of memory will be needed.
1.5.6
Conditional Execution
This is an interesting feature of ARM instruction set. While most of the existing architecures
allow only branches to be executed conditionally, ARM allows all instructions to be executed
conditionally. The most significant four bits of each instruction are reserved to hold the 16
condition codes. There are four flags N, Z, C, and V. The following are the interpretaions of
these four bit settings.
0000
: EQ (Equal)
0001
0010
0011
0100
0101
0110
0111
1000
1001
1010
1011
1100
1110
1111
:
:
:
:
:
:
:
:
:
:
:
:
:
:
23
NE (Not equal)
HS/CS (Unsigned higher or same)
LO/CC (Unsigned lower)
MI (Negative)
PL (Positive or zero)
VS (Overflow)
VC (No overflow)
HI (Unsigned higher)
LS (Set unsigned lower or same)
GE (Greater or equal)
LT (Lower)
GT (Greater)
AL (Always)
NV (Reserved)
We have already seen how to set the conditional flags in instruction execution. For example,
the use of mnemonic ADDS in place of ADD sets the flags, whereas, the basic ADD instruction
does not affect any of the flags. Now, to execute an instruction conditionally, we have to use
mnemonis like EQADD. For example,
EQADD R0, R1, R2
EQADDS R1, R2, R3, LSL #2
1.5.7
Branch instruction
24
31
28 27
Condition
25 24 23
1 0 1
Offset
1.5.8
Swap instruction
A careful look into the ARM instructions will reveal that a single instruction can read/write
at the most one memory location. The swap instriction is an exception to this. Swap is an
atomic operation in which a memory read is followed by a memory write which moves byte or
word between registers and memory. Its format is,
SWP Rd, Rm, [Rn]
SWPB Rd, Rm, [Rn]
The execution proceeds as follows (as depicted in Fig 1.8).
1. It is a two-cycle operation.
2. Content of memory location pointed to by register Rn is copied into a temorary space.
3. Content of register Rm is copied into the memory location.
4. Content of the temporary space is copied into the register Rd.
25
Thus, to effect an interchange between the registers Rd and Rm, they should be made the
same register. In that case, content of the memory location is swapped with the register in
an atomic operation. This can be exploited to implement the semaphore operations of the
operating system.
1
Rn
temp
3
2
Rm
Rd
Memory
Figure 1.8: Swap instruction execution
1.5.9
The status registers (that is, CPSR and SPSRs) can only be modified indirectly. The instruction MSR moves content from CPSR/SPSR to selected general purpose register (GPR),
whereas, the MRS moves content of selected GPR to CPSR/SPSR. The instruction can only
be executed in the privileged modes.
1.6
THUMB instructions
As we have noted earlier, the ARM processor supports two different instruction sets ARM
and THUMB. While the ARM instruction set has got 32-bit instructions, THUMB instructions
are 16-bit in length. These instructions are stored in a compressed form. The instructions are
decompressed into ARM instructions and then executed by the processor. Fig 1.9 shows
the THUMB instruction processing. Fig 1.10 shows the THUMB instruction decompressor.
THUMB instruction set must always be entered by runing a BX/BLX instruction. These
26
THUMB
Decompressor
Instruction
Pipeline
16 bit
THUMB code
ARM
Instruction
Decoder
To B Operand Bus
data in
immediate field
select ARM or
THUMB stream
Mux
Thumb decompressor
Mux
Instruction pipeline
select high or
low halfword
27
CPSR
28
1.7
Exceptions in ARM
The exceptions can be raised in ARM from different situations. These can be split into three
different categories.
1. Exceptions can occur through execution of instructions. These include the software
interrupts, undefined instructions and memory abort.
29
2. Exceptions can also be caused as a side-effect of an instruction, like data fetch failure.
3. Other sources like RESET, FIQ, and IRQ interrupts.
For all these exceptions, the mechanism followed is similar. The processor switches to the
privileged mode, current value of PC(R15) + 4 is saved into the link register (R14) of the
privileged mode. Current CPSR is saved in the SPSR of the privileged mode. The IRQ
interrupts are disabled. If the exception raised is an FIQ, the FIQ interrupts are also disabled.
Next, the program counter is loaded with the exception vector address, and the execution of
the exception starts. Table 1.3 shows the vector table for different ARM exceptions. Two
Table 1.3: ARM7 vector table
Exception type
Mode
Reset
Supervisor
Undefined
Undefined instruction
Software interrupt (SWI)
Supervisor
Prefetch abort (instruction fetch abort) Abort
Data Abort (data access memory abort) Abort
IRQ
IRQ (interrupt)
FIQ (fast interrupt)
FIQ
Meaning
0x00000000
0x00000004
0x00000008
0x0000000C
0x00000010
0x00000018
0x0000001C
1.8
Programming examples
In this section, we will look into a few programming examples of moderate complexity. The
idea is to get an understanding about how to use the ARM instructions in an assembly language
program.
30
1.8.1
Assume that the register R0 holds the start address of the data block (each data is 4 bytes
long) and R2 holds the number of elements in the block. The maximum of the elements will
be stored in register R1.
EOR
CMP
BEQ
R1, R1, R1
R2, #0
Over
LDR
CMP
BCC
MOV
R3, [R0]
R3, R1
Looptest
R1, R3
ADD
SUBS
BNE
R0, R0, #4
R2, R2, #1
Loop
;increment pointer R0
;decrement number of elements left
;if not done, loop
Loop
Looptest
Over
;R1 holds the largest
1.8.2
In this example we assume that the two strings are pointed to by the registers R0 and R1
respectively, both the strings are null terminated (that is, the last byte is 0), and that the
result will be stored in register R2. If the two strings match exactly, the content of R2 will be
0, else it will be -1.
Loop
LDRB
LDRB
CMP
BNE
CMP
BEQ
ADD
ADD
BAL
R3, [R0]
R4, [R1]
R3, R4
Notsame
R3, #0
Same
R0, R0, #1
R1, R1, #1
Loop
1.9. CONCLUSION
31
Notsame
MOV
R2, #-1
BAL Over
MOV
;mark matched
Same
R2, #0
Over
;R2 holds the match
1.9
Conclusion
In this chapter, we have seen a broad overview of ARM microcontroller architecture. It has
many interesting architectural features combining both CISC and RISC philosophy. The facilities provided in FIQ handling reduces response time to critical interrupts. The controlled
modification of status bits by arithmetic instructions and conditional instruction execution are
unique to the processor. The large number of system calls that can be supported by SWI
instruction is definitely helpful for the OS designers. Avaialability of large number of registers
help in speedy context switch. The highly compact and low-power THUMB instructions make
the processor ideal choice for small, low-cost embedded system development.
Exercises
1.1 How does a typical microcontroller differ from a microprocessor? What are the typical
architecural blocks present in a microcontroller?
1.2 What are the factors guiding the choice of a microcontroller for an embedded application?
1.3 What is meant by application specific ARM processor? How does it help the embedded
system designer?
1.4 How does in-circuit emulation helps in system development?
1.5 Explain the internal structure of ARM processor.
1.6 How does the ALE signal of ARM differ from ALE of x86 family of processors?
1.7 Explain the functional diagram of ARM processor.
1.8 Enumerate the evolution of various pipelining structures in ARM.
32
1.9 Compare and contrast between RISC and CISC instructions. Explain how in ARM both
of the philosophies have been merged.
1.10 How many general and special-purpose registers are there in ARM? Explain their functionality.
1.11 State the role of R13, R14, R15. How does R15 differ from program counter in general
CPU?
1.12 Enumerate various operating modes of ARM.
1.13 Mention different data types available in ARM. How does a big-endian type differ from
little-endian type?
1.14 Suppose the number 0x12345678 is stored at memory location 1000 in big-endian format.
If the processor now assumes the data to be in little-endian format, what will it get if it
reads (a) a byte, (b) a half-word, (c) a word from location 1000? In the previous example,
suppose that the data bus lines are connected to the memory in the reverse order, that
is bit-0 from processor is connected to data line 31 of memory, 1 to 30, and so on. What
value will be read in caeses (a), (b) and (c) noted above?
1.15 Enumerate the types of instructions in ARM.
1.16 Why is it the case that all possible 32-bit operands cannot be specified as immediate
operands in ARM data processing instructions? Out of the following numbers, which
can be represented: 56, 1000, 513, -1?
1.17 How to set condition flags in arithmetic operations in ARM? What advantage does this
preferential setting provide us?
1.18 Mention different ways in which the address of a memory operand can be specified in
ARM.
1.19 Why do you think that not providing separate stack-space in ARM may be beneficial?
Show the content of the stack after executing each of these instructions on a partially
filled stack STMFD SP!, {R0-R12}, STMED SP!, {R0-R12}, STMFA SP!, {R0-R12},
STMEA SP!, {R0-R12}.
1.20 What do you mean by early-termination? If the operands for multiplication are 12 and
55, which one should be used as the source operand?
1.21 State the differences between the ways in which software interrupts are handled in Intel
processors and in ARM.
1.9. CONCLUSION
33
1.22 Explain the conditional execution of statements in ARM. For the following piece of highlevel code, show the corresponding assembly-level code for ARM. Assume all the variables
are available in registers.
if a > b then
x=y+z
else
x=yz
1.23 What are the advantages of program counter relative jump over absolute jump? Why
do you think that the 24-bit offset in the branch instruction structure is sufficient
for 32MB branch range in ARM? Assuming that the current program counter is at
0x10000000, compute the 24-bit offset for the addresses 0x10000020 and 0x0FFFFFFC.
1.24 How does BX/BLX differ from B/BL?
1.25 What is meant by an atomic operation? How is the SWAP intruction different from others
from atomicity viewpoint? How does it help in implementing semaphores, explain with
suitable example.
1.26 What is THUMB? How does the THUMB instruction set differ from ARM instruction
set? Are the THUMB instructions executed directly?
1.27 Why are the FIQ interrupts can be handled faster than IRQ interrupts in ARM?
1.28 Write the following programs in the assembly language of ARM.
(a) Sort a given set of numbers.
(b) Find second highest of a set of numbers without actually sorting it.
(c) Concatenate two null terminated strings.
(d) Look for a specific substring within a null terminated string.
1.29 Perform a comparative study between the microcontrollers (i) ARM, (ii) 8051, (iii) PIC,
(iv) AVR.