| 1 | .. _tcg-ops-ref: |
| 2 | |
| 3 | ******************************* |
| 4 | TCG Intermediate Representation |
| 5 | ******************************* |
| 6 | |
| 7 | Introduction |
| 8 | ============ |
| 9 | |
| 10 | TCG (Tiny Code Generator) began as a generic backend for a C compiler. |
| 11 | It was simplified to be used in QEMU. It also has its roots in the |
| 12 | QOP code generator written by Paul Brook. |
| 13 | |
| 14 | Definitions |
| 15 | =========== |
| 16 | |
| 17 | The TCG *target* is the architecture for which we generate the code. |
| 18 | It is of course not the same as the "target" of QEMU which is the |
| 19 | emulated architecture. As TCG started as a generic C backend used |
| 20 | for cross compiling, the assumption was that TCG target might be |
| 21 | different from the host, although this is never the case for QEMU. |
| 22 | |
| 23 | In this document, we use *guest* to specify what architecture we are |
| 24 | emulating; *target* always means the TCG target, the machine on which |
| 25 | we are running QEMU. |
| 26 | |
| 27 | An operation with *undefined behavior* may result in a crash. |
| 28 | |
| 29 | An operation with *unspecified behavior* shall not crash. However, |
| 30 | the result may be one of several possibilities so may be considered |
| 31 | an *undefined result*. |
| 32 | |
| 33 | Basic Blocks |
| 34 | ============ |
| 35 | |
| 36 | A TCG *basic block* is a single entry, multiple exit region which |
| 37 | corresponds to a list of instructions terminated by a label, or |
| 38 | any branch instruction. |
| 39 | |
| 40 | A TCG *extended basic block* is a single entry, multiple exit region |
| 41 | which corresponds to a list of instructions terminated by a label or |
| 42 | an unconditional branch. Specifically, an extended basic block is |
| 43 | a sequence of basic blocks connected by the fall-through paths of |
| 44 | zero or more conditional branch instructions. |
| 45 | |
| 46 | Operations |
| 47 | ========== |
| 48 | |
| 49 | TCG instructions or *ops* operate on TCG *variables*, both of which |
| 50 | are strongly typed. Each instruction has a fixed number of output |
| 51 | variable operands, input variable operands and constant operands. |
| 52 | Vector instructions have a field specifying the element size within |
| 53 | the vector. The notable exception is the call instruction which has |
| 54 | a variable number of outputs and inputs. |
| 55 | |
| 56 | In the textual form, output operands usually come first, followed by |
| 57 | input operands, followed by constant operands. The output type is |
| 58 | included in the instruction name. Constants are prefixed with a '$'. |
| 59 | |
| 60 | .. code-block:: none |
| 61 | |
| 62 | add_i32 t0, t1, t2 /* (t0 <- t1 + t2) */ |
| 63 | |
| 64 | Variables |
| 65 | ========= |
| 66 | |
| 67 | * ``TEMP_FIXED`` |
| 68 | |
| 69 | There is one TCG *fixed global* variable, ``cpu_env``, which is |
| 70 | live in all translation blocks, and holds a pointer to ``CPUArchState``. |
| 71 | This variable is held in a host cpu register at all times in all |
| 72 | translation blocks. |
| 73 | |
| 74 | * ``TEMP_GLOBAL`` |
| 75 | |
| 76 | A TCG *global* is a variable which is live in all translation blocks, |
| 77 | and corresponds to memory location that is within ``CPUArchState``. |
| 78 | These may be specified as an offset from ``cpu_env``, in which case |
| 79 | they are called *direct globals*, or may be specified as an offset |
| 80 | from a direct global, in which case they are called *indirect globals*. |
| 81 | Even indirect globals should still reference memory within |
| 82 | ``CPUArchState``. All TCG globals are defined during |
| 83 | ``TCGCPUOps.initialize``, before any translation blocks are generated. |
| 84 | |
| 85 | * ``TEMP_CONST`` |
| 86 | |
| 87 | A TCG *constant* is a variable which is live throughout the entire |
| 88 | translation block, and contains a constant value. These variables |
| 89 | are allocated on demand during translation and are hashed so that |
| 90 | there is exactly one variable holding a given value. |
| 91 | |
| 92 | * ``TEMP_TB`` |
| 93 | |
| 94 | A TCG *translation block temporary* is a variable which is live |
| 95 | throughout the entire translation block, but dies on any exit. |
| 96 | These temporaries are allocated explicitly during translation. |
| 97 | |
| 98 | * ``TEMP_EBB`` |
| 99 | |
| 100 | A TCG *extended basic block temporary* is a variable which is live |
| 101 | throughout an extended basic block, but dies on any exit. |
| 102 | These temporaries are allocated explicitly during translation. |
| 103 | |
| 104 | Types |
| 105 | ===== |
| 106 | |
| 107 | * ``TCG_TYPE_I32`` |
| 108 | |
| 109 | A 32-bit integer. |
| 110 | |
| 111 | * ``TCG_TYPE_I64`` |
| 112 | |
| 113 | A 64-bit integer. For 32-bit hosts, such variables are split into a pair |
| 114 | of variables with ``type=TCG_TYPE_I32`` and ``base_type=TCG_TYPE_I64``. |
| 115 | The ``temp_subindex`` for each indicates where it falls within the |
| 116 | host-endian representation. |
| 117 | |
| 118 | * ``TCG_TYPE_PTR`` |
| 119 | |
| 120 | An alias for ``TCG_TYPE_I32`` or ``TCG_TYPE_I64``, depending on the size |
| 121 | of a pointer for the host. |
| 122 | |
| 123 | * ``TCG_TYPE_REG`` |
| 124 | |
| 125 | An alias for ``TCG_TYPE_I32`` or ``TCG_TYPE_I64``, depending on the size |
| 126 | of the integer registers for the host. This may be larger |
| 127 | than ``TCG_TYPE_PTR`` depending on the host ABI. |
| 128 | |
| 129 | * ``TCG_TYPE_I128`` |
| 130 | |
| 131 | A 128-bit integer. For all hosts, such variables are split into a number |
| 132 | of variables with ``type=TCG_TYPE_REG`` and ``base_type=TCG_TYPE_I128``. |
| 133 | The ``temp_subindex`` for each indicates where it falls within the |
| 134 | host-endian representation. |
| 135 | |
| 136 | * ``TCG_TYPE_V64`` |
| 137 | |
| 138 | A 64-bit vector. This type is valid only if the TCG target |
| 139 | sets ``TCG_TARGET_HAS_v64``. |
| 140 | |
| 141 | * ``TCG_TYPE_V128`` |
| 142 | |
| 143 | A 128-bit vector. This type is valid only if the TCG target |
| 144 | sets ``TCG_TARGET_HAS_v128``. |
| 145 | |
| 146 | * ``TCG_TYPE_V256`` |
| 147 | |
| 148 | A 256-bit vector. This type is valid only if the TCG target |
| 149 | sets ``TCG_TARGET_HAS_v256``. |
| 150 | |
| 151 | Helpers |
| 152 | ======= |
| 153 | |
| 154 | Helpers are registered in a guest-specific ``helper.h``, |
| 155 | which is processed to generate ``tcg_gen_helper_*`` functions. |
| 156 | With these functions it is possible to call a function taking |
| 157 | i32, i64, i128 or pointer types. |
| 158 | |
| 159 | By default, before calling a helper, all globals are stored at their |
| 160 | canonical location. By default, the helper is allowed to modify the |
| 161 | CPU state (including the state represented by tcg globals) |
| 162 | or may raise an exception. This default can be overridden using the |
| 163 | following function modifiers: |
| 164 | |
| 165 | * ``TCG_CALL_NO_WRITE_GLOBALS`` |
| 166 | |
| 167 | The helper does not modify any globals, but may read them. |
| 168 | Globals will be saved to their canonical location before calling helpers, |
| 169 | but need not be reloaded afterwards. |
| 170 | |
| 171 | * ``TCG_CALL_NO_READ_GLOBALS`` |
| 172 | |
| 173 | The helper does not read globals, either directly or via an exception. |
| 174 | They will not be saved to their canonical locations before calling |
| 175 | the helper. This implies ``TCG_CALL_NO_WRITE_GLOBALS``. |
| 176 | |
| 177 | * ``TCG_CALL_NO_SIDE_EFFECTS`` |
| 178 | |
| 179 | The call to the helper function may be removed if the return value is |
| 180 | not used. This means that it may not modify any CPU state nor may it |
| 181 | raise an exception. |
| 182 | |
| 183 | Code Optimizations |
| 184 | ================== |
| 185 | |
| 186 | When generating instructions, you can count on at least the following |
| 187 | optimizations: |
| 188 | |
| 189 | - Single instructions are simplified, e.g. |
| 190 | |
| 191 | .. code-block:: none |
| 192 | |
| 193 | and_i32 t0, t0, $0xffffffff |
| 194 | |
| 195 | is suppressed. |
| 196 | |
| 197 | - A liveness analysis is done at the basic block level. The |
| 198 | information is used to suppress moves from a dead variable to |
| 199 | another one. It is also used to remove instructions which compute |
| 200 | dead results. The later is especially useful for condition code |
| 201 | optimization in QEMU. |
| 202 | |
| 203 | In the following example: |
| 204 | |
| 205 | .. code-block:: none |
| 206 | |
| 207 | add_i32 t0, t1, t2 |
| 208 | add_i32 t0, t0, $1 |
| 209 | mov_i32 t0, $1 |
| 210 | |
| 211 | only the last instruction is kept. |
| 212 | |
| 213 | |
| 214 | Instruction Reference |
| 215 | ===================== |
| 216 | |
| 217 | Function call |
| 218 | ------------- |
| 219 | |
| 220 | .. list-table:: |
| 221 | |
| 222 | * - call *<ret>* *<params>* ptr |
| 223 | |
| 224 | - | call function 'ptr' (pointer type) |
| 225 | | |
| 226 | | *<ret>* optional 32 bit or 64 bit return value |
| 227 | | *<params>* optional 32 bit or 64 bit parameters |
| 228 | |
| 229 | Jumps/Labels |
| 230 | ------------ |
| 231 | |
| 232 | .. list-table:: |
| 233 | |
| 234 | * - set_label $label |
| 235 | |
| 236 | - | Define label 'label' at the current program point. |
| 237 | |
| 238 | * - br $label |
| 239 | |
| 240 | - | Jump to label. |
| 241 | |
| 242 | * - brcond *t0*, *t1*, *cond*, *label* |
| 243 | |
| 244 | - | Conditional jump if *t0* *cond* *t1* is true. *cond* can be: |
| 245 | | |
| 246 | | ``TCG_COND_EQ`` |
| 247 | | ``TCG_COND_NE`` |
| 248 | | ``TCG_COND_LT /* signed */`` |
| 249 | | ``TCG_COND_GE /* signed */`` |
| 250 | | ``TCG_COND_LE /* signed */`` |
| 251 | | ``TCG_COND_GT /* signed */`` |
| 252 | | ``TCG_COND_LTU /* unsigned */`` |
| 253 | | ``TCG_COND_GEU /* unsigned */`` |
| 254 | | ``TCG_COND_LEU /* unsigned */`` |
| 255 | | ``TCG_COND_GTU /* unsigned */`` |
| 256 | | ``TCG_COND_TSTEQ /* t1 & t2 == 0 */`` |
| 257 | | ``TCG_COND_TSTNE /* t1 & t2 != 0 */`` |
| 258 | |
| 259 | Arithmetic |
| 260 | ---------- |
| 261 | |
| 262 | .. list-table:: |
| 263 | |
| 264 | * - add *t0*, *t1*, *t2* |
| 265 | |
| 266 | - | *t0* = *t1* + *t2* |
| 267 | |
| 268 | * - sub *t0*, *t1*, *t2* |
| 269 | |
| 270 | - | *t0* = *t1* - *t2* |
| 271 | |
| 272 | * - neg *t0*, *t1* |
| 273 | |
| 274 | - | *t0* = -*t1* (two's complement) |
| 275 | |
| 276 | * - mul *t0*, *t1*, *t2* |
| 277 | |
| 278 | - | *t0* = *t1* * *t2* |
| 279 | |
| 280 | * - divs *t0*, *t1*, *t2* |
| 281 | |
| 282 | - | *t0* = *t1* / *t2* (signed) |
| 283 | | Undefined behavior if division by zero or overflow. |
| 284 | |
| 285 | * - divu *t0*, *t1*, *t2* |
| 286 | |
| 287 | - | *t0* = *t1* / *t2* (unsigned) |
| 288 | | Undefined behavior if division by zero. |
| 289 | |
| 290 | * - rems *t0*, *t1*, *t2* |
| 291 | |
| 292 | - | *t0* = *t1* % *t2* (signed) |
| 293 | | Undefined behavior if division by zero or overflow. |
| 294 | |
| 295 | * - remu *t0*, *t1*, *t2* |
| 296 | |
| 297 | - | *t0* = *t1* % *t2* (unsigned) |
| 298 | | Undefined behavior if division by zero. |
| 299 | |
| 300 | * - divs2 *q*, *r*, *nl*, *nh*, *d* |
| 301 | |
| 302 | - | *q* = *nh:nl* / *d* (signed) |
| 303 | | *r* = *nh:nl* % *d* |
| 304 | | Undefined behaviour if division by zero, or the double-word |
| 305 | numerator divided by the single-word divisor does not fit |
| 306 | within the single-word quotient. The code generator will |
| 307 | pass *nh* as a simple sign-extension of *nl*, so the only |
| 308 | overflow should be *INT_MIN* / -1. |
| 309 | |
| 310 | * - divu2 *q*, *r*, *nl*, *nh*, *d* |
| 311 | |
| 312 | - | *q* = *nh:nl* / *d* (unsigned) |
| 313 | | *r* = *nh:nl* % *d* |
| 314 | | Undefined behaviour if division by zero, or the double-word |
| 315 | numerator divided by the single-word divisor does not fit |
| 316 | within the single-word quotient. The code generator will |
| 317 | pass 0 to *nh* to make a simple zero-extension of *nl*, |
| 318 | so overflow should never occur. |
| 319 | |
| 320 | * - smax *t0*, *t1*, *t2* |
| 321 | |
| 322 | umax *t0*, *t1*, *t2* |
| 323 | |
| 324 | - | *t0* = MAX(*t1*, *t2*), for signed and unsigned integers. |
| 325 | |
| 326 | * - smin *t0*, *t1*, *t2* |
| 327 | |
| 328 | umin *t0*, *t1*, *t2* |
| 329 | |
| 330 | - | *t0* = MIN(*t1*, *t2*), for signed and unsigned integers. |
| 331 | |
| 332 | Logical |
| 333 | ------- |
| 334 | |
| 335 | .. list-table:: |
| 336 | |
| 337 | * - and *t0*, *t1*, *t2* |
| 338 | |
| 339 | - | *t0* = *t1* & *t2* |
| 340 | |
| 341 | * - or *t0*, *t1*, *t2* |
| 342 | |
| 343 | - | *t0* = *t1* | *t2* |
| 344 | |
| 345 | * - xor *t0*, *t1*, *t2* |
| 346 | |
| 347 | - | *t0* = *t1* ^ *t2* |
| 348 | |
| 349 | * - not *t0*, *t1* |
| 350 | |
| 351 | - | *t0* = ~\ *t1* |
| 352 | |
| 353 | * - andc *t0*, *t1*, *t2* |
| 354 | |
| 355 | - | *t0* = *t1* & ~\ *t2* |
| 356 | |
| 357 | * - eqv *t0*, *t1*, *t2* |
| 358 | |
| 359 | - | *t0* = ~(*t1* ^ *t2*), or equivalently, *t0* = *t1* ^ ~\ *t2* |
| 360 | |
| 361 | * - nand *t0*, *t1*, *t2* |
| 362 | |
| 363 | - | *t0* = ~(*t1* & *t2*) |
| 364 | |
| 365 | * - nor *t0*, *t1*, *t2* |
| 366 | |
| 367 | - | *t0* = ~(*t1* | *t2*) |
| 368 | |
| 369 | * - orc *t0*, *t1*, *t2* |
| 370 | |
| 371 | - | *t0* = *t1* | ~\ *t2* |
| 372 | |
| 373 | * - clz *t0*, *t1*, *t2* |
| 374 | |
| 375 | - | *t0* = *t1* ? clz(*t1*) : *t2* |
| 376 | |
| 377 | * - ctz *t0*, *t1*, *t2* |
| 378 | |
| 379 | - | *t0* = *t1* ? ctz(*t1*) : *t2* |
| 380 | |
| 381 | * - ctpop *t0*, *t1* |
| 382 | |
| 383 | - | *t0* = number of bits set in *t1* |
| 384 | | |
| 385 | | The name *ctpop* is short for "count population", and matches |
| 386 | the function name used in ``include/qemu/host-utils.h``. |
| 387 | |
| 388 | |
| 389 | Shifts/Rotates |
| 390 | -------------- |
| 391 | |
| 392 | .. list-table:: |
| 393 | |
| 394 | * - shl *t0*, *t1*, *t2* |
| 395 | |
| 396 | - | *t0* = *t1* << *t2* |
| 397 | | Unspecified behavior for negative or out-of-range shifts. |
| 398 | |
| 399 | * - shr *t0*, *t1*, *t2* |
| 400 | |
| 401 | - | *t0* = *t1* >> *t2* (unsigned) |
| 402 | | Unspecified behavior for negative or out-of-range shifts. |
| 403 | |
| 404 | * - sar *t0*, *t1*, *t2* |
| 405 | |
| 406 | - | *t0* = *t1* >> *t2* (signed) |
| 407 | | Unspecified behavior for negative or out-of-range shifts. |
| 408 | |
| 409 | * - rotl *t0*, *t1*, *t2* |
| 410 | |
| 411 | - | Rotation of *t2* bits to the left |
| 412 | | Unspecified behavior for negative or out-of-range shifts. |
| 413 | |
| 414 | * - rotr *t0*, *t1*, *t2* |
| 415 | |
| 416 | - | Rotation of *t2* bits to the right. |
| 417 | | Unspecified behavior for negative or out-of-range shifts. |
| 418 | |
| 419 | |
| 420 | Misc |
| 421 | ---- |
| 422 | |
| 423 | .. list-table:: |
| 424 | |
| 425 | * - mov *t0*, *t1* |
| 426 | |
| 427 | - | *t0* = *t1* |
| 428 | | Move *t1* to *t0*. |
| 429 | |
| 430 | * - bswap16 *t0*, *t1*, *flags* |
| 431 | |
| 432 | - | 16 bit byte swap on the low bits of a 32/64 bit input. |
| 433 | | |
| 434 | | If *flags* & ``TCG_BSWAP_IZ``, then *t1* is known to be zero-extended from bit 15. |
| 435 | | If *flags* & ``TCG_BSWAP_OZ``, then *t0* will be zero-extended from bit 15. |
| 436 | | If *flags* & ``TCG_BSWAP_OS``, then *t0* will be sign-extended from bit 15. |
| 437 | | |
| 438 | | If neither ``TCG_BSWAP_OZ`` nor ``TCG_BSWAP_OS`` are set, then the bits of *t0* above bit 15 may contain any value. |
| 439 | |
| 440 | * - bswap32 *t0*, *t1*, *flags* |
| 441 | |
| 442 | - | 32 bit byte swap. The flags are the same as for bswap16, except |
| 443 | they apply from bit 31 instead of bit 15. On TCG_TYPE_I32, the |
| 444 | flags should be zero. |
| 445 | |
| 446 | * - bswap64 *t0*, *t1*, *flags* |
| 447 | |
| 448 | - | 64 bit byte swap. The flags are ignored, but still present |
| 449 | for consistency with the other bswap opcodes. For future |
| 450 | compatibility, the flags should be zero. |
| 451 | |
| 452 | * - discard_i32/i64 *t0* |
| 453 | |
| 454 | - | Indicate that the value of *t0* won't be used later. It is useful to |
| 455 | force dead code elimination. |
| 456 | |
| 457 | * - deposit *dest*, *t1*, *t2*, *pos*, *len* |
| 458 | |
| 459 | - | Deposit *t2* as a bitfield into *t1*, placing the result in *dest*. |
| 460 | | |
| 461 | | The bitfield is described by *pos*/*len*, which are immediate values: |
| 462 | | |
| 463 | | *len* - the length of the bitfield |
| 464 | | *pos* - the position of the first bit, counting from the LSB |
| 465 | | |
| 466 | | For example, "deposit dest, t1, t2, 8, 4" indicates a 4-bit field |
| 467 | at bit 8. This operation would be equivalent to |
| 468 | | |
| 469 | | *dest* = (*t1* & ~0x0f00) | ((*t2* << 8) & 0x0f00) |
| 470 | | |
| 471 | | on TCG_TYPE_I32. |
| 472 | |
| 473 | * - extract *dest*, *t1*, *pos*, *len* |
| 474 | |
| 475 | sextract *dest*, *t1*, *pos*, *len* |
| 476 | |
| 477 | - | Extract a bitfield from *t1*, placing the result in *dest*. |
| 478 | | |
| 479 | | The bitfield is described by *pos*/*len*, which are immediate values, |
| 480 | as above for deposit. For extract_*, the result will be extended |
| 481 | to the left with zeros; for sextract_*, the result will be extended |
| 482 | to the left with copies of the bitfield sign bit at *pos* + *len* - 1. |
| 483 | | |
| 484 | | For example, "sextract dest, t1, 8, 4" indicates a 4-bit field |
| 485 | at bit 8. This operation would be equivalent to |
| 486 | | |
| 487 | | *dest* = (*t1* << 20) >> 28 |
| 488 | | |
| 489 | | (using an arithmetic right shift) on TCG_TYPE_I32. |
| 490 | |
| 491 | * - extract2 *dest*, *t1*, *t2*, *pos* |
| 492 | |
| 493 | - | For TCG_TYPE_I{N}, extract an N-bit quantity from the concatenation |
| 494 | of *t2*:*t1*, beginning at *pos*. The tcg_gen_extract2_{i32,i64} expander |
| 495 | accepts 0 <= *pos* <= N as inputs. The backend code generator will |
| 496 | not see either 0 or N as inputs for these opcodes. |
| 497 | |
| 498 | * - extrl_i64_i32 *t0*, *t1* |
| 499 | |
| 500 | - | For 64-bit hosts only, extract the low 32-bits of input *t1* and place it |
| 501 | into 32-bit output *t0*. Depending on the host, this may be a simple move, |
| 502 | or may require additional canonicalization. |
| 503 | |
| 504 | * - extrh_i64_i32 *t0*, *t1* |
| 505 | |
| 506 | - | For 64-bit hosts only, extract the high 32-bits of input *t1* and place it |
| 507 | into 32-bit output *t0*. Depending on the host, this may be a simple shift, |
| 508 | or may require additional canonicalization. |
| 509 | |
| 510 | * - revbit8 *dest*, *t1* |
| 511 | |
| 512 | - | Reverse the 8 bits within each byte of input *t1* with |
| 513 | | output in *dest*; the byte order is unchanged. |
| 514 | |
| 515 | * - revbit32 *dest*, *t1*, *flags* |
| 516 | |
| 517 | - | Reverse the 32 bits of the lower 32 bits of input *t1* |
| 518 | | with output in *dest*. On TCG_TYPE_I64, *flags* control |
| 519 | | any required sign or zero extension of the result in |
| 520 | | the same way as for bswap32. |
| 521 | | On TCG_TYPE_I32, *flags* should be zero. |
| 522 | |
| 523 | * - revbit64 *dest*, *t1* |
| 524 | |
| 525 | - | Reverse the 64 bits of input *t1* with output in *dest*. |
| 526 | |
| 527 | Conditional moves |
| 528 | ----------------- |
| 529 | |
| 530 | .. list-table:: |
| 531 | |
| 532 | * - setcond *dest*, *t1*, *t2*, *cond* |
| 533 | |
| 534 | - | *dest* = (*t1* *cond* *t2*) |
| 535 | | |
| 536 | | Set *dest* to 1 if (*t1* *cond* *t2*) is true, otherwise set to 0. |
| 537 | |
| 538 | * - negsetcond *dest*, *t1*, *t2*, *cond* |
| 539 | |
| 540 | - | *dest* = -(*t1* *cond* *t2*) |
| 541 | | |
| 542 | | Set *dest* to -1 if (*t1* *cond* *t2*) is true, otherwise set to 0. |
| 543 | |
| 544 | * - movcond *dest*, *c1*, *c2*, *v1*, *v2*, *cond* |
| 545 | |
| 546 | - | *dest* = (*c1* *cond* *c2* ? *v1* : *v2*) |
| 547 | | |
| 548 | | Set *dest* to *v1* if (*c1* *cond* *c2*) is true, otherwise set to *v2*. |
| 549 | |
| 550 | |
| 551 | Type conversions |
| 552 | ---------------- |
| 553 | |
| 554 | .. list-table:: |
| 555 | |
| 556 | * - ext_i32_i64 *t0*, *t1* |
| 557 | |
| 558 | - | Convert *t1* (32 bit) to *t0* (64 bit) and does sign extension |
| 559 | |
| 560 | * - extu_i32_i64 *t0*, *t1* |
| 561 | |
| 562 | - | Convert *t1* (32 bit) to *t0* (64 bit) and does zero extension |
| 563 | |
| 564 | * - trunc_i64_i32 *t0*, *t1* |
| 565 | |
| 566 | - | Truncate *t1* (64 bit) to *t0* (32 bit) |
| 567 | |
| 568 | * - concat_i32_i64 *t0*, *t1*, *t2* |
| 569 | |
| 570 | - | Construct *t0* (64-bit) taking the low half from *t1* (32 bit) and the high half |
| 571 | from *t2* (32 bit). |
| 572 | |
| 573 | * - concat32_i64 *t0*, *t1*, *t2* |
| 574 | |
| 575 | - | Construct *t0* (64-bit) taking the low half from *t1* (64 bit) and the high half |
| 576 | from *t2* (64 bit). |
| 577 | |
| 578 | |
| 579 | Load/Store |
| 580 | ---------- |
| 581 | |
| 582 | .. list-table:: |
| 583 | |
| 584 | * - ld_i32/i64 *t0*, *t1*, *offset* |
| 585 | |
| 586 | ld8s_i32/i64 *t0*, *t1*, *offset* |
| 587 | |
| 588 | ld8u_i32/i64 *t0*, *t1*, *offset* |
| 589 | |
| 590 | ld16s_i32/i64 *t0*, *t1*, *offset* |
| 591 | |
| 592 | ld16u_i32/i64 *t0*, *t1*, *offset* |
| 593 | |
| 594 | ld32s_i64 *t0*, *t1*, *offset* |
| 595 | |
| 596 | ld32u_i64 *t0*, *t1*, *offset* |
| 597 | |
| 598 | - | *t0* = read(*t1* + *offset*) |
| 599 | | |
| 600 | | Load 8, 16, 32 or 64 bits with or without sign extension from host memory. |
| 601 | *offset* must be a constant. |
| 602 | |
| 603 | * - st_i32/i64 *t0*, *t1*, *offset* |
| 604 | |
| 605 | st8_i32/i64 *t0*, *t1*, *offset* |
| 606 | |
| 607 | st16_i32/i64 *t0*, *t1*, *offset* |
| 608 | |
| 609 | st32_i64 *t0*, *t1*, *offset* |
| 610 | |
| 611 | - | write(*t0*, *t1* + *offset*) |
| 612 | | |
| 613 | | Write 8, 16, 32 or 64 bits to host memory. |
| 614 | |
| 615 | All this opcodes assume that the pointed host memory doesn't correspond |
| 616 | to a global. In the latter case the behaviour is unpredictable. |
| 617 | |
| 618 | |
| 619 | Multiword arithmetic support |
| 620 | ---------------------------- |
| 621 | |
| 622 | .. list-table:: |
| 623 | |
| 624 | * - addco *t0*, *t1*, *t2* |
| 625 | |
| 626 | - | Compute *t0* = *t1* + *t2* and in addition output to the |
| 627 | carry bit provided by the host architecture. |
| 628 | |
| 629 | * - addci *t0*, *t1*, *t2* |
| 630 | |
| 631 | - | Compute *t0* = *t1* + *t2* + *C*, where *C* is the |
| 632 | input carry bit provided by the host architecture. |
| 633 | The output carry bit need not be computed. |
| 634 | |
| 635 | * - addcio *t0*, *t1*, *t2* |
| 636 | |
| 637 | - | Compute *t0* = *t1* + *t2* + *C*, where *C* is the |
| 638 | input carry bit provided by the host architecture, |
| 639 | and also compute the output carry bit. |
| 640 | |
| 641 | * - addc1o *t0*, *t1*, *t2* |
| 642 | |
| 643 | - | Compute *t0* = *t1* + *t2* + 1, and in addition output to the |
| 644 | carry bit provided by the host architecture. This is akin to |
| 645 | *addcio* with a fixed carry-in value of 1. |
| 646 | | This is intended to be used by the optimization pass, |
| 647 | intermediate to complete folding of the addition chain. |
| 648 | In some cases complete folding is not possible and this |
| 649 | opcode will remain until output. If this happens, the |
| 650 | code generator will use ``tcg_out_set_carry`` and then |
| 651 | the output routine for *addcio*. |
| 652 | |
| 653 | * - subbo *t0*, *t1*, *t2* |
| 654 | |
| 655 | - | Compute *t0* = *t1* - *t2* and in addition output to the |
| 656 | borrow bit provided by the host architecture. |
| 657 | | Depending on the host architecture, the carry bit may or may not be |
| 658 | identical to the borrow bit. Thus the addc\* and subb\* |
| 659 | opcodes must not be mixed. |
| 660 | |
| 661 | * - subbi *t0*, *t1*, *t2* |
| 662 | |
| 663 | - | Compute *t0* = *t1* - *t2* - *B*, where *B* is the |
| 664 | input borrow bit provided by the host architecture. |
| 665 | The output borrow bit need not be computed. |
| 666 | |
| 667 | * - subbio *t0*, *t1*, *t2* |
| 668 | |
| 669 | - | Compute *t0* = *t1* - *t2* - *B*, where *B* is the |
| 670 | input borrow bit provided by the host architecture, |
| 671 | and also compute the output borrow bit. |
| 672 | |
| 673 | * - subb1o *t0*, *t1*, *t2* |
| 674 | |
| 675 | - | Compute *t0* = *t1* - *t2* - 1, and in addition output to the |
| 676 | borrow bit provided by the host architecture. This is akin to |
| 677 | *subbio* with a fixed borrow-in value of 1. |
| 678 | | This is intended to be used by the optimization pass, |
| 679 | intermediate to complete folding of the subtraction chain. |
| 680 | In some cases complete folding is not possible and this |
| 681 | opcode will remain until output. If this happens, the |
| 682 | code generator will use ``tcg_out_set_borrow`` and then |
| 683 | the output routine for *subbio*. |
| 684 | |
| 685 | * - mulu2 *t0_low*, *t0_high*, *t1*, *t2* |
| 686 | |
| 687 | - | Similar to mul, except two unsigned inputs *t1* and *t2* yielding the full |
| 688 | double-word product *t0*. The latter is returned in two single-word outputs. |
| 689 | |
| 690 | * - muls2 *t0_low*, *t0_high*, *t1*, *t2* |
| 691 | |
| 692 | - | Similar to mulu2, except the two inputs *t1* and *t2* are signed. |
| 693 | |
| 694 | * - mulsh *t0*, *t1*, *t2* |
| 695 | |
| 696 | muluh *t0*, *t1*, *t2* |
| 697 | |
| 698 | - | Provide the high part of a signed or unsigned multiply, respectively. |
| 699 | | |
| 700 | | If mulu2/muls2 are not provided by the backend, the tcg-op generator |
| 701 | can obtain the same results by emitting a pair of opcodes, mul + muluh/mulsh. |
| 702 | |
| 703 | |
| 704 | Memory Barrier support |
| 705 | ---------------------- |
| 706 | |
| 707 | .. list-table:: |
| 708 | |
| 709 | * - mb *<$arg>* |
| 710 | |
| 711 | - | Generate a target memory barrier instruction to ensure memory ordering |
| 712 | as being enforced by a corresponding guest memory barrier instruction. |
| 713 | | |
| 714 | | The ordering enforced by the backend may be stricter than the ordering |
| 715 | required by the guest. It cannot be weaker. This opcode takes a constant |
| 716 | argument which is required to generate the appropriate barrier |
| 717 | instruction. The backend should take care to emit the target barrier |
| 718 | instruction only when necessary i.e., for SMP guests and when MTTCG is |
| 719 | enabled. |
| 720 | | |
| 721 | | The guest translators should generate this opcode for all guest instructions |
| 722 | which have ordering side effects. |
| 723 | | |
| 724 | | Please see :ref:`atomics-ref` for more information on memory barriers. |
| 725 | |
| 726 | |
| 727 | QEMU specific operations |
| 728 | ------------------------ |
| 729 | |
| 730 | .. list-table:: |
| 731 | |
| 732 | * - exit_tb *t0* |
| 733 | |
| 734 | - | Exit the current TB and return the value *t0* (word type). |
| 735 | |
| 736 | * - goto_tb *index* |
| 737 | |
| 738 | - | Exit the current TB and jump to the TB index *index* (constant) if the |
| 739 | current TB was linked to this TB. Otherwise execute the next |
| 740 | instructions. Only indices 0 and 1 are valid and tcg_gen_goto_tb may be issued |
| 741 | at most once with each slot index per TB. |
| 742 | |
| 743 | * - lookup_and_goto_ptr *tb_addr* |
| 744 | |
| 745 | - | Look up a TB address *tb_addr* and jump to it if valid. If not valid, |
| 746 | jump to the TCG epilogue to go back to the exec loop. |
| 747 | | |
| 748 | | This operation is optional. If the TCG backend does not implement the |
| 749 | goto_ptr opcode, emitting this op is equivalent to emitting exit_tb(0). |
| 750 | |
| 751 | * - qemu_ld_i32/i64/i128 *t0*, *t1*, *flags*, *memidx* |
| 752 | |
| 753 | qemu_st_i32/i64/i128 *t0*, *t1*, *flags*, *memidx* |
| 754 | |
| 755 | - | Load data at the guest address *t1* into *t0*, or store data in *t0* at guest |
| 756 | address *t1*. The _i32/_i64/_i128 size applies to the size of the input/output |
| 757 | register *t0* only. The address *t1* is always sized according to the guest, |
| 758 | and the width of the memory operation is controlled by *flags*. |
| 759 | | |
| 760 | | Both *t0* and *t1* may be split into little-endian ordered pairs of registers |
| 761 | if dealing with 64-bit quantities on a 32-bit host, or 128-bit quantities on |
| 762 | a 64-bit host. |
| 763 | | |
| 764 | | The *memidx* selects the qemu tlb index to use (e.g. user or kernel access). |
| 765 | The flags are the MemOp bits, selecting the sign, width, and endianness |
| 766 | of the memory access. |
| 767 | | |
| 768 | | For a 32-bit host, qemu_ld/st_i64 is guaranteed to only be used with a |
| 769 | 64-bit memory access specified in *flags*. |
| 770 | | |
| 771 | | For qemu_ld/st_i128, these are only supported for a 64-bit host. |
| 772 | |
| 773 | |
| 774 | Host vector operations |
| 775 | ---------------------- |
| 776 | |
| 777 | All of the vector ops have two parameters, ``TCGOP_TYPE`` & ``TCGOP_VECE``. |
| 778 | The former specifies the length of the vector as a TCGType; the latter |
| 779 | specifies the length of the element (if applicable) in log2 8-bit units. |
| 780 | |
| 781 | .. list-table:: |
| 782 | |
| 783 | * - mov_vec *v0*, *v1* |
| 784 | |
| 785 | ld_vec *v0*, *t1* |
| 786 | |
| 787 | st_vec *v0*, *t1* |
| 788 | |
| 789 | - | Move, load and store. |
| 790 | |
| 791 | * - dup_vec *v0*, *r1* |
| 792 | |
| 793 | - | Duplicate the low N bits of *r1* into TYPE/VECE copies across *v0*. |
| 794 | |
| 795 | * - dupi_vec *v0*, *c* |
| 796 | |
| 797 | - | Similarly, for a constant. |
| 798 | | Smaller values will be replicated to host register size by the expanders. |
| 799 | |
| 800 | * - add_vec *v0*, *v1*, *v2* |
| 801 | |
| 802 | - | *v0* = *v1* + *v2*, in elements across the vector. |
| 803 | |
| 804 | * - sub_vec *v0*, *v1*, *v2* |
| 805 | |
| 806 | - | Similarly, *v0* = *v1* - *v2*. |
| 807 | |
| 808 | * - mul_vec *v0*, *v1*, *v2* |
| 809 | |
| 810 | - | Similarly, *v0* = *v1* * *v2*. |
| 811 | |
| 812 | * - neg_vec *v0*, *v1* |
| 813 | |
| 814 | - | Similarly, *v0* = -*v1*. |
| 815 | |
| 816 | * - abs_vec *v0*, *v1* |
| 817 | |
| 818 | - | Similarly, *v0* = *v1* < 0 ? -*v1* : *v1*, in elements across the vector. |
| 819 | |
| 820 | * - smin_vec *v0*, *v1*, *v2* |
| 821 | |
| 822 | umin_vec *v0*, *v1*, *v2* |
| 823 | |
| 824 | - | Similarly, *v0* = MIN(*v1*, *v2*), for signed and unsigned element types. |
| 825 | |
| 826 | * - smax_vec *v0*, *v1*, *v2* |
| 827 | |
| 828 | umax_vec *v0*, *v1*, *v2* |
| 829 | |
| 830 | - | Similarly, *v0* = MAX(*v1*, *v2*), for signed and unsigned element types. |
| 831 | |
| 832 | * - ssadd_vec *v0*, *v1*, *v2* |
| 833 | |
| 834 | sssub_vec *v0*, *v1*, *v2* |
| 835 | |
| 836 | usadd_vec *v0*, *v1*, *v2* |
| 837 | |
| 838 | ussub_vec *v0*, *v1*, *v2* |
| 839 | |
| 840 | - | Signed and unsigned saturating addition and subtraction. |
| 841 | | |
| 842 | | If the true result is not representable within the element type, the |
| 843 | element is set to the minimum or maximum value for the type. |
| 844 | |
| 845 | * - and_vec *v0*, *v1*, *v2* |
| 846 | |
| 847 | nand_vec *v0*, *v1*, *v2* |
| 848 | |
| 849 | or_vec *v0*, *v1*, *v2* |
| 850 | |
| 851 | nor_vec *v0*, *v1*, *v2* |
| 852 | |
| 853 | xor_vec *v0*, *v1*, *v2* |
| 854 | |
| 855 | eqv_vec *v0*, *v1*, *v2* |
| 856 | |
| 857 | andc_vec *v0*, *v1*, *v2* |
| 858 | |
| 859 | orc_vec *v0*, *v1*, *v2* |
| 860 | |
| 861 | not_vec *v0*, *v1* |
| 862 | |
| 863 | - | Similarly, logical operations with and without complement. |
| 864 | | |
| 865 | | Note that VECE is unused. |
| 866 | |
| 867 | * - shli_vec *v0*, *v1*, *i2* |
| 868 | |
| 869 | shls_vec *v0*, *v1*, *s2* |
| 870 | |
| 871 | - | Shift all elements from v1 by a scalar *i2*/*s2*. I.e. |
| 872 | |
| 873 | .. code-block:: c |
| 874 | |
| 875 | for (i = 0; i < TYPE/VECE; ++i) { |
| 876 | v0[i] = v1[i] << s2; |
| 877 | } |
| 878 | |
| 879 | * - shri_vec *v0*, *v1*, *i2* |
| 880 | |
| 881 | sari_vec *v0*, *v1*, *i2* |
| 882 | |
| 883 | rotli_vec *v0*, *v1*, *i2* |
| 884 | |
| 885 | shrs_vec *v0*, *v1*, *s2* |
| 886 | |
| 887 | sars_vec *v0*, *v1*, *s2* |
| 888 | |
| 889 | rotls_vec *v0*, *v1*, *s2* |
| 890 | |
| 891 | - | Similarly for logical and arithmetic right shift, and left rotate. |
| 892 | |
| 893 | * - shlv_vec *v0*, *v1*, *v2* |
| 894 | |
| 895 | - | Shift elements from *v1* by elements from *v2*. I.e. |
| 896 | |
| 897 | .. code-block:: c |
| 898 | |
| 899 | for (i = 0; i < TYPE/VECE; ++i) { |
| 900 | v0[i] = v1[i] << v2[i]; |
| 901 | } |
| 902 | |
| 903 | * - shrv_vec *v0*, *v1*, *v2* |
| 904 | |
| 905 | sarv_vec *v0*, *v1*, *v2* |
| 906 | |
| 907 | rotlv_vec *v0*, *v1*, *v2* |
| 908 | |
| 909 | rotrv_vec *v0*, *v1*, *v2* |
| 910 | |
| 911 | - | Similarly for logical and arithmetic right shift, and rotates. |
| 912 | |
| 913 | * - cmp_vec *v0*, *v1*, *v2*, *cond* |
| 914 | |
| 915 | - | Compare vectors by element, storing -1 for true and 0 for false. |
| 916 | |
| 917 | * - bitsel_vec *v0*, *v1*, *v2*, *v3* |
| 918 | |
| 919 | - | Bitwise select, *v0* = (*v2* & *v1*) | (*v3* & ~\ *v1*), across the entire vector. |
| 920 | |
| 921 | * - cmpsel_vec *v0*, *c1*, *c2*, *v3*, *v4*, *cond* |
| 922 | |
| 923 | - | Select elements based on comparison results: |
| 924 | |
| 925 | .. code-block:: c |
| 926 | |
| 927 | for (i = 0; i < n; ++i) { |
| 928 | v0[i] = (c1[i] cond c2[i]) ? v3[i] : v4[i]. |
| 929 | } |
| 930 | |
| 931 | **Note 1**: Some shortcuts are defined when the last operand is known to be |
| 932 | a constant (e.g. addi for add, movi for mov). |
| 933 | |
| 934 | **Note 2**: When using TCG, the opcodes must never be generated directly |
| 935 | as some of them may not be available as "real" opcodes. Always use the |
| 936 | function tcg_gen_xxx(args). |
| 937 | |
| 938 | |
| 939 | Backend |
| 940 | ======= |
| 941 | |
| 942 | ``tcg-target.h`` contains the target specific definitions. ``tcg-target.c.inc`` |
| 943 | contains the target specific code; it is #included by ``tcg/tcg.c``, rather |
| 944 | than being a standalone C file. |
| 945 | |
| 946 | Assumptions |
| 947 | ----------- |
| 948 | |
| 949 | The target word size (``TCG_TARGET_REG_BITS``) is expected to be 64 bit. |
| 950 | It is expected that the pointer has the same size as the word. |
| 951 | |
| 952 | Values are transferred between 32 and 64-bit registers using the |
| 953 | following ops: |
| 954 | |
| 955 | - extrl_i64_i32 |
| 956 | - extrh_i64_i32 |
| 957 | - ext_i32_i64 |
| 958 | - extu_i32_i64 |
| 959 | |
| 960 | They ensure that the values are correctly truncated or extended when |
| 961 | moved from a 32-bit to a 64-bit register or vice-versa. Note that the |
| 962 | extrl_i64_i32 and extrh_i64_i32 are optional ops. It is not necessary |
| 963 | to implement them if all the following conditions are met: |
| 964 | |
| 965 | - 64-bit registers can hold 32-bit values |
| 966 | - 32-bit values in a 64-bit register do not need to stay zero or |
| 967 | sign extended |
| 968 | - all 32-bit TCG ops ignore the high part of 64-bit registers |
| 969 | |
| 970 | Floating point operations are not supported in this version. A |
| 971 | previous incarnation of the code generator had full support of them, |
| 972 | but it is better to concentrate on integer operations first. |
| 973 | |
| 974 | Constraints |
| 975 | ---------------- |
| 976 | |
| 977 | GCC like constraints are used to define the constraints of every |
| 978 | instruction. Memory constraints are not supported in this |
| 979 | version. Aliases are specified in the input operands as for GCC. |
| 980 | |
| 981 | The same register may be used for both an input and an output, even when |
| 982 | they are not explicitly aliased. If an op expands to multiple target |
| 983 | instructions then care must be taken to avoid clobbering input values. |
| 984 | GCC style "early clobber" outputs are supported, with '``&``'. |
| 985 | |
| 986 | A target can define specific register or constant constraints. If an |
| 987 | operation uses a constant input constraint which does not allow all |
| 988 | constants, it must also accept registers in order to have a fallback. |
| 989 | The constraint '``i``' is defined generically to accept any constant. |
| 990 | The constraint '``r``' is not defined generically, but is consistently |
| 991 | used by each backend to indicate all registers. If ``TCG_REG_ZERO`` |
| 992 | is defined by the backend, the constraint '``z``' is defined generically |
| 993 | to map constant 0 to the hardware zero register. |
| 994 | |
| 995 | The movi_i32 and movi_i64 operations must accept any constants. |
| 996 | |
| 997 | The mov_i32 and mov_i64 operations must accept any registers of the |
| 998 | same type. |
| 999 | |
| 1000 | The ld/st/sti instructions must accept signed 32 bit constant offsets. |
| 1001 | This can be implemented by reserving a specific register in which to |
| 1002 | compute the address if the offset is too big. |
| 1003 | |
| 1004 | The ld/st instructions must accept any destination (ld) or source (st) |
| 1005 | register. |
| 1006 | |
| 1007 | The sti instruction may fail if it cannot store the given constant. |
| 1008 | |
| 1009 | Function call assumptions |
| 1010 | ------------------------- |
| 1011 | |
| 1012 | - The only supported types for parameters and return value are: 32 and |
| 1013 | 64 bit integers and pointer. |
| 1014 | - The stack grows downwards. |
| 1015 | - The first N parameters are passed in registers. |
| 1016 | - The next parameters are passed on the stack by storing them as words. |
| 1017 | - Some registers are clobbered during the call. |
| 1018 | - The function can return 0 or 1 value in registers. On a 32 bit |
| 1019 | target, functions must be able to return 2 values in registers for |
| 1020 | 64 bit return type. |
| 1021 | |
| 1022 | |
| 1023 | Recommended coding rules for best performance |
| 1024 | ============================================= |
| 1025 | |
| 1026 | - Use globals to represent the parts of the QEMU CPU state which are |
| 1027 | often modified, e.g. the integer registers and the condition |
| 1028 | codes. TCG will be able to use host registers to store them. |
| 1029 | |
| 1030 | - Don't hesitate to use helpers for complicated or seldom used guest |
| 1031 | instructions. There is little performance advantage in using TCG to |
| 1032 | implement guest instructions taking more than about twenty TCG |
| 1033 | instructions. Note that this rule of thumb is more applicable to |
| 1034 | helpers doing complex logic or arithmetic, where the C compiler has |
| 1035 | scope to do a good job of optimisation; it is less relevant where |
| 1036 | the instruction is mostly doing loads and stores, and in those cases |
| 1037 | inline TCG may still be faster for longer sequences. |
| 1038 | |
| 1039 | - Use the 'discard' instruction if you know that TCG won't be able to |
| 1040 | prove that a given global is "dead" at a given program point. The |
| 1041 | x86 guest uses it to improve the condition codes optimisation. |